# NFL_ROBBED_METHODOLOGY.md — "which NFL team gets robbed the most?"

The research, the model and the sensitivity analysis behind the `NFLRobbed`
episode. Written 2026-09-16. **The code is the specification**:
`scripts/nfl_robbed.py` is the only place anything is computed, and this document
explains it rather than restating it.

**Rebuild everything:**

```bash
../.venv-nba/bin/python scripts/nfl_robbed.py     # ~6 min incl. 3 permutation runs
node scripts/nfl_robbed_scene.mjs --print
../.venv/bin/python3 scripts/vo_nflrobbed.py
```

⚠️ `.venv-nba` is this repo's **parquet venv** (pandas + pyarrow + numpy), created
for the NBA episodes. This script installs nothing and reuses it. `pyarrow` must
never be installed into the shared `.venv`.

---

## 1 · What is actually being claimed

**Tennessee has the best case in the league that the officiating has gone against
them. That is not the same as proof, and the episode says so on screen and on the
tape.**

The precise claim: over the five completed seasons 2021–2025, on the penalties
that require a referee to exercise judgment, the Titans are **last of 32** at
**−3.3 points of win probability per game**, which is **−0.56 wins per 17-game
season** — while being **9th of 32** on the penalties a camera settles.

What it is **not**: a claim that any individual call was wrong. See §6.

---

## 2 · The window: 2021–2025, and why five seasons

Every figure is measured over the **five completed regular seasons 2021–2025** —
1,359 games, 16,303 accepted penalties.

Long enough that each club has ~85 games and ~500 flags called against it, which
is what the permutation test needs to have any power at all. Short enough that it
is a claim about the team that exists now rather than about a coaching staff two
regimes ago. Postseason is excluded: the sample is small, the crews are selected
on merit rather than assigned in rotation, and mixing the two would confound the
crew analysis in §5.

Relocations are folded into the current franchise (OAK→LV, SD→LAC, STL→LA), the
same convention `NFL_FAN_VALUE_METHODOLOGY.md` uses.

---

## 3 · Move one: the unit is win probability, not yards

**Every published version of this question counts flags or counts yards.** Both
are averages over a distribution that is not remotely symmetric.

`nflfastR` ships a win-probability model and reports `wpa` — win probability
added for the possession team — on every play, including penalty plays. Flipping
the sign for the penalised side gives the **win probability that flag cost the
team it was called on**. Across all 16,303:

| | win probability points |
|---|---|
| median flag | **1.84** |
| mean | 2.65 |
| 90th percentile | 5.90 |
| 99th percentile | **14.16** |
| worst single flag, five seasons | **66.54** |

The worst **1%** of flags carry **8.3%** of all the win probability penalties move
in five seasons; the worst 10% carry **36.2%**. A holding call on 3rd-and-8 in a
tie game late is simply not the same object as a false start on the opening
drive, and "10 yards" says it is.

This move alone changes the answer: net penalty yards says **Chicago**
(−10.0 yd/game); win probability says the **Jets** (−0.72 wins/season). The two
rankings correlate at r = 0.77 — related, not the same.

⚠️ **This inherits nflfastR's win-probability model wholesale**, including its
assumptions and its error. It is the best public instrument available and it is
not ours; a different WP model would move these numbers somewhat, though the
facts/opinions *contrast* in §4 is a within-model comparison and is far more
robust than the levels.

---

## 4 · Move two: facts against opinions

**This is the episode's one editorial idea and the only place a judgment call of
our own enters the model.**

Every accepted penalty is coded into one of two classes:

- **MECHANICAL — a fact.** Position or timing, settled by the tape. He was across
  the line or he was not; the play clock hit zero or it did not. A crew has no
  discretion about the *degree*. False start, delay of game, offside, neutral
  zone infraction, encroachment, illegal formation/shift/motion, too many men,
  ineligible downfield, the kickoff placement fouls. **6,854 flags.**
- **JUDGMENT — an opinion.** Somebody decided that a grab materially restricted,
  that contact was unnecessary, that a hit was forcible, that there was no
  receiver in the area. Offensive and defensive holding, pass interference both
  ways, illegal contact, illegal use of hands, unnecessary roughness, roughing
  the passer, face mask, the block fouls, intentional grounding, unsportsmanlike
  conduct, taunting, tripping. **9,449 flags — 58% of all penalties.**

**Only the second kind can be wrong in a team's favour or against it.**

### 4.1 The result that licenses everything after it

Across the 32 clubs, the mechanical differential and the judgment differential
correlate at **r = +0.015**.

**How disciplined a team is tells you essentially nothing about how the judgment
calls go for it.** That is what makes it legitimate to look at the judgment
column on its own — if the two moved together, a judgment deficit would just be
another way of measuring the same indiscipline, and there would be no episode.

### 4.2 And it disqualifies the win-probability leader

The Jets are worst in the league on the whole penalty differential — and they are
**31st of 32 on the mechanical flags**. They commit them. They are not being
robbed; they are undisciplined, and the naive win-probability ranking is largely
measuring that.

Tennessee is the opposite shape: **9th of 32 on the facts, last of 32 on the
opinions.**

### 4.3 The coding is ours, so it is stress-tested five ways

`sensitivity()` re-runs the entire ranking under five codings — the base, face
mask moved to facts, intentional grounding moved to opinions, ineligible
downfield moved to opinions, and formation/shift/motion moved to opinions.

**Tennessee is worst under all five.** The build fails if that ever stops being
true (`assert all(s["worst"] == worst for s in sens)`), because the episode names
a club and may not name one it cannot defend.

---

## 5 · Move three: the null, and what the referees themselves do

### 5.1 The permutation test

The null keeps **every judgment flag the league actually threw** — its own win
probability cost, its own season, its own penalty type — and reshuffles only
**which team it was called on**, within (season × penalty type). 4,000 draws.

Reshuffling within season-and-type is the point. A null that pooled everything
would let a 2021 roughing-the-passer cost stand in for a 2025 illegal contact and
would make every club look extreme.

**Tennessee: z = −2.27, p = 0.015.**

### 5.2 ⚠️ And the multiple-comparisons correction, which is load-bearing

**32 clubs were tested. Somebody has to finish last.**

Taking the minimum across all 32 clubs in each of the 4,000 draws, a deficit at
least this large turns up in **43%** of them.

So the honest statement, and the one the episode makes: *Tennessee has the
strongest case in the league, it survives every version of the coding, and it is
still not proof.* Both numbers are on screen and both are spoken. An episode that
reported only the 0.015 would be overclaiming, and this is exactly the caveat the
playbook says separates the channel from slop.

### 5.3 Do the crews themselves show up?

Using `games.csv`'s `referee` column, one row per game, residualised on both
clubs' identities, against a null that reshuffles crew assignments within season:

| question | crew-to-crew sd | p |
|---|---|---|
| does a crew favour the **home team**? | 0.0158 | **0.41** — no |
| does a crew differ in **how much a game swings** on flags? | 0.0314 | **0.019** — yes |

**Crews are not biased toward the home side.** They do differ in severity: which
crew you draw genuinely changes how much of the game turns on penalties. That is
real, and it is symmetric — a loose crew flags both teams — so it is variance, not
robbery.

⚠️ **THE BUG THAT MADE THIS READ "NO EFFECT" AT FIRST, and it is a trap anybody
re-deriving this will hit.** A team's penalty differential is *exactly* the
negative of its opponent's. So in a per-team-game fit, the two rows of every game
cancel and **every crew's residual mean is zero by construction** — the test
reports "referees explain nothing" no matter what the data says, with no error and
no warning. The unit has to be the GAME, with a signed home-team quantity.

---

## 6 · ⚠️ What this cannot see

**There is no public record of which NFL calls were correct.** The league grades
its officials privately and publishes nothing equivalent to the NBA's Last Two
Minute report.

So this measures **how the judgment calls fell**, never whether any one of them
was right. A club can finish last here because it is genuinely being
mis-officiated, or because it plays a style that draws more judgment flags, or
because of five seasons of bad luck. The permutation test addresses the third and
says it is unlikely on its own and unremarkable across 32 clubs; it cannot
separate the first two at all.

That limitation is stated on screen in beat 5 and spoken in the closing line, and
it is why the claim is "the best case in the league" rather than "the refs are
doing it".

---

## 7 · Where everything lives

| path | what |
|---|---|
| `scripts/nfl_robbed.py` | **the model.** Run with `../.venv-nba/bin/python` |
| `scripts/nfl_robbed_scene.mjs` | selects what the episode may say → `public/nfl/robbed_scene.json` |
| `scripts/vo_nflrobbed.py` | the tape. Run with `../.venv/bin/python3` |
| `data/nfl/nfl_robbed.csv` | **the shippable dataset**, 32 rows |
| `public/nfl/robbed.json` | the model's full output incl. permutation + sensitivity |
| `public/nfl/robbed_scene.json` | everything the episode is allowed to say |
| `src/series/NFLRobbed.tsx` | the composition — beat table, VO mounting, watermark |
| `src/series/nflrobbed/` | the six beats + the shell |
| `.cache/nfl/` | nflverse play-by-play parquet, ~100MB, gitignored |

Shared with the fan-value episode and **not** rebuilt here: `public/nfl/logos/`,
`public/nfl/colors.json`, and the club lat/lon in `public/nfl/raw.json`.

**Sources.** nflverse play-by-play (`nflverse-data`, pbp release) for every play
and the win-probability model; `nflverse/nfldata` `games.csv` for the referee
assignment. Both keyless and public.
