reelgorithm.py

The most penalised team is not the most robbed team

We took all 16,303 accepted penalties from five NFL seasons, priced every one of them in win probability instead of yards, and then split them into the calls a camera settles and the calls a referee decides. The team that loses the most is the team that commits the most. The team with the best case is ninth in the league at breaking the rules.

“We got robbed” is the oldest argument in football, and it is almost always made with the wrong number. Somebody posts net penalty yards, somebody else posts flag counts, and everyone goes home certain.

Both of those numbers have the same two problems. They treat every flag as the same size, and they treat every flag as the same kind. Fix those two things and the answer changes twice.

Move one: a flag is not ten yards

Penalty yardage is an average over a distribution that is nothing like symmetric. A false start on the opening drive and a defensive holding call on third down in a tie game with four minutes left are not the same object, and “five yards” versus “five yards” says they are.

So we priced each flag the way the game itself does: in win probability. The nflfastR model that ships with the nflverse play-by-play reports win probability added on every play, penalty plays included. Flip the sign for the side the flag was called on and you have what that call cost the team it was called on.

what one accepted flag costs, 16,303 of them, 2021–2025
1.84
median flag
win probability points
5.90
90th percentile
14.16
99th percentile
66.54
worst single flag
in five seasons

The worst 1% of flags carry 8.3% of all the win probability that penalties move in five seasons. The worst 10% carry 36.2%. A ranking built on yards is spending most of its resolution on calls that barely mattered.

And the switch moves the answer immediately. On net penalty yards the most-wronged team in football is Chicago, at −10.04 yards a game. On win probability it is the New York Jets, at −0.72 wins per season. The two rankings correlate at r = 0.773 — related, and not the same.

Move two: only half of a flag is an opinion

This is the one editorial judgement in the whole model, and it is the reason the answer is not simply “the Jets”.

Some penalties are facts. He was across the line or he was not. The play clock hit zero or it did not. A crew has no discretion about the degree of a false start. Others are opinions — somebody decided that a grab materially restricted, that contact was unnecessary, that there was no receiver in the area.

classwhat settles itexamplesflags
Mechanicalthe tapefalse start, delay of game, offside, encroachment, illegal formation / shift / motion, too many men, ineligible downfield6,854
Judgmenta refereeoffensive and defensive holding, pass interference, illegal contact, unnecessary roughness, roughing the passer, face mask, illegal blocks, grounding, unsportsmanlike conduct9,449
58% of all accepted penalties are judgment calls. Only that half can be wrong in a team’s favour or against it.

Splitting them is only legitimate if the two halves are actually different things. They are, and this is the result that licenses everything after it:

Across the 32 clubs, the mechanical differential and the judgment differential correlate at r = +0.015.

How disciplined a team is tells you essentially nothing about how the judgment calls go for it. If those two columns moved together, a judgment deficit would just be another way of measuring indiscipline and there would be nothing here. They do not, so the opinion column can be read on its own.

Which disqualifies the leader

The Jets top the win probability ladder. They are also 31st of 32 at committing the flags a camera settles.

cluball flags, wins/seasonrank on the factsrank on the opinions
Jets−0.7231st of 3231st of 32
Titans−0.489th of 3232nd of 32
Ravens−0.4219th28th
Patriots−0.3823rd27th
Saints−0.3710th30th
49ers−0.3513th29th
Rank 1 is the cleanest record on the facts and the most favourable treatment on the opinions. The Jets are not being robbed. They are undisciplined, and the naive ladder is largely measuring that.

Tennessee is the opposite shape. Ninth of thirty-two on the facts. Last of thirty-two on the opinions. They are a well-drilled football team, by the measure that cannot be argued with, and the calls that require somebody to make a decision have gone against them more than against anyone else in the league.

The size of it

tennessee titans — judgment calls only, 85 games, 2021–2025
−3.31
win probability points
per game
−0.56
wins per
17-game season
−2.27
z against
the permutation null
9th
of 32 on the flags
a camera settles

Half a win a season is not a conspiracy and it is not nothing. It is roughly the difference between a nine-win team and an eight-win team, every year, for five years, arriving entirely through calls where a human being had to decide.

And it is not spread evenly across the rulebook. Sorting Tennessee’s judgment deficit by the kind of flag, one category is nearly half of it:

judgment categorypercentile in the leaguewin probability / game
Pass interference16th−1.60
Roughness19th−0.81
Coverage19th−0.57
Offensive holding41st−0.28
Illegal blocks19th−0.23
Conduct & grounding53rd+0.16
The six categories partition the 9,449 judgment flags exactly — no flag is in two and none is in none. Pass interference alone is 48% of Tennessee’s whole deficit.

Which is the category with the largest yardage swing and the least reviewable definition in the sport. That is not proof of anything. It is, however, exactly where you would look first.

Do the referees themselves show up?

The obvious next question is whether particular crews explain it. The games file names the referee, so this is answerable: 1,358 games, 20 crews, residualised on both clubs’ identities, against a null that reshuffles crew assignments within season.

questioncrew-to-crew sdpanswer
Does a crew favour the home team?0.01580.41No
Do crews differ in how much a game swings on flags?0.03140.019Yes
Crews are not biased toward the home side. They do differ in severity — which crew you draw genuinely changes how much of the game turns on penalties.

That second result is real and it is symmetric. A loose crew flags both teams. It is variance in the experience of watching a game, not a thumb on a scale.

the bug that nearly buried this

The first version of this test reported “referees explain nothing” with a clean p-value, and it was wrong by construction. A team’s penalty differential is exactly the negative of its opponent’s, so in a per-team-game fit the two rows of every game cancel and every crew’s residual mean is zero no matter what the data says.

No error, no warning, a plausible number. The unit has to be the game, with a signed home-team quantity. Anyone re-deriving this will hit the same wall.

The part that argues against us

Two things, and the second is the bigger one.

The coding is ours

Which penalties count as opinions is a decision we made, so the whole ranking was re-run under five different codings — face mask moved to facts, intentional grounding moved to opinions, ineligible downfield moved to opinions, formation and shift and motion moved to opinions, and the published one.

codingworst clubrunner-up
PublishedTitansJets
Face mask counted as a factTitansJets
Intentional grounding counted as an opinionTitansJets
Ineligible downfield counted as an opinionTitansJets
Formation / shift / motion counted as opinionsTitansJets
Tennessee is last under all five. The build fails if that ever stops being true — the episode names a club and may not name one it cannot defend.

Somebody has to finish last

The permutation test keeps every judgment flag the league actually threw — its own win probability cost, its own season, its own penalty type — and reshuffles only which team it was called on, within season and type. Four thousand draws.

Tennessee’s deficit comes out at p = 0.015. Taken on its own that is a result.

It was not taken on its own. Thirty-two clubs were tested. Taking the minimum across all 32 in each of the 4,000 draws, a deficit at least this large turns up in 43% of them.

Tennessee has the strongest case in the league, it survives every version of the coding, and it is still not proof.

Both numbers are in the video and both are in the file. An analysis that reported only the 0.015 would be overclaiming, and the gap between those two p-values is most of what separates a real finding from a viral one.

and the thing this cannot see at all

There is no public record of which NFL calls were correct. The league grades its officials privately and publishes nothing equivalent to the NBA’s Last Two Minute report.

So this measures how the judgment calls fell, never whether any one of them was right. A club can finish last here because it is genuinely being mis-officiated, because it plays a style that draws more judgment flags, or because of five seasons of bad luck. The permutation test addresses the third. It cannot separate the first two at all, and no public dataset can.

The other end of the table

For completeness, because a deficit only means something against a surplus: the club the judgment calls have favoured most over five seasons is Minnesota, at +4.35 win probability points a game, or +0.74 wins a season. The Giants are second at +0.34, Houston third at +0.33 — and Houston is 30th of 32 at committing the mechanical fouls, which is the mirror image of the Titans and just as strange.

The full 32-club table, both columns, is in the CSV.

Get the data

Every club, both differentials, the permutation z and p, the win equivalents and the game counts — the file everything above is computed from.

the data — everything above is reproducible from these

nfl-robbed-2026.csv — 32 clubs × 18 columns: mechanical and judgment differentials in flags, win probability and yards, the combined figures, games played, the permutation z and one-club p, and the win-per-season equivalents for both the judgment column and all penalties.
nfl-robbed-methodology.md — the window and why it is five seasons, the full mechanical / judgment coding, the permutation design, the multiple-comparisons correction, the crew analysis and the zero-sum bug inside it, and what the model cannot see.

Sources: nflverse play-by-play (nflverse-data, pbp release), regular seasons 2021–2025 — 1,359 games and 16,303 accepted penalties; the nflfastR win-probability model (Baldwin) that ships with it; and nflverse/nfldata games.csv for referee assignment. Both keyless and public. Relocated franchises are folded into the current club (OAK→LV, SD→LAC, STL→LA). Postseason is excluded: the sample is small and the crews are selected on merit rather than assigned in rotation, which would confound the crew analysis. The win-probability model is inherited wholesale, including its assumptions and its error — a different model would move the levels, though the facts-against-opinions contrast is a within-model comparison and is far more robust than the levels are.