# The Fifth — dataset & methodology

*865 athlete-years · 1,394 event-years retrieved · nine gates, one of which failed
and rewrote a claim.*

**`phelps-world-number-ones.csv`** (built as `phelps_field.csv`) — every
athlete-year, 1985–2025, in which a swimmer
finished a calendar year ranked **#1 in the world** in at least one individual
long-course event, with how many such events and how many stroke families they
spanned. Produced by `data/phelps_field.py` from the World Aquatics official
season world rankings.

| column | meaning |
|---|---|
| `sex` | `M` / `F` — rankings are separate competitions and are never pooled into one comparison |
| `year` | calendar year of the ranking |
| `athlete` | name as published by World Aquatics |
| `country` | national federation code |
| `world_titles` | number of individual events ranked #1 in the world that year |
| `stroke_families` | how many of free / back / breast / fly / IM those events span |
| `events` | the events themselves, pipe-separated |
| `archive_depth_ok` | 1 if every counted event-year cleared the coverage floor |

## What is measured

For each sex, each of the 17 individual long-course events, and each year
1985–2025, the official season world ranking is retrieved in `BEST_TIMES` mode
— one row per swimmer, their fastest legal time of that year. The swimmer
holding rank 1 is that event-year's world #1. Athlete-years are then formed by
grouping those winners by swimmer and year.

Relays are excluded. A relay leg is not an individual world ranking, and
including them would inflate the count for exactly the swimmers the video is
about.

## The thing that makes this hard, and why there is no σ anywhere

The obvious version of this analysis computes how many standard deviations the
subject sits above the field and converts that into a "one in N" via a normal
tail. **This dataset cannot support that, and neither could the ones that get
used for it.** Two independent reasons:

1. **Season-best times are not normal in the fast tail.** They are bounded below
   by physiology and thinned by selection. A Gaussian extrapolated five or more
   σ past the mean is wrong by orders of magnitude, in a direction nobody can
   bound from the data.

2. **Archive depth is digitisation, not participation.** The identical query
   returns **6** ranked swimmers for the 1990 men's 200m butterfly and **4,794**
   for 2025. The sport did not get 800× deeper. Any z-score computed against
   "the field" is therefore mostly a measurement of how much of that era has
   been entered into the database.

So rarity in this dataset is only ever an **observed frequency** — a count of
athlete-years that actually happened. Nothing is extrapolated past the data.

## Gates

`data/phelps_field.py` refuses to write output unless all nine pass.

- **G0 — known answer.** Michael Phelps's 2008 must come back as world #1 in
  exactly the five events the historical record says: 200 free, 100 fly, 200
  fly, 200 IM, 400 IM. This validates the parser against a fact established
  independently of the pipeline, before any claim is derived from it.
- **G1 — coverage floor.** An event-year counts only if the archive holds at
  least 25 ranked swimmers, so "world #1" is never an artefact of a six-row
  table. Sparse early event-years are dropped, not silently counted.
- **G2 — supersuit independence.** 2008–09 was the polyurethane era, and the
  subject's headline year sits inside it. The analysis is recomputed with both
  years deleted entirely; the record must still stand. A finding that only
  exists in the suit years is a finding about suits.
- **G3 — rarity is a count.** The reported rarity must equal an observed number
  of athlete-years. Asserted structurally so no future edit can quietly swap in
  a fitted tail.
- **G4 — threshold robustness.** The whole analysis is re-run at a *top-3*
  cutoff instead of *#1*. The subject must still hold the maximum. A result that
  exists only at one arbitrary cutoff is a cutoff, not a result.
- **G5 — breadth.** The subject's record year must span at least three stroke
  families, and no other athlete-year that also spans three or more may hold as
  many titles. *This gate was rewritten during the build.* Its first version
  asserted the subject spanned strictly **more** families than anyone, and it
  failed — Lochte 2011, Marchand 2024, Otto 1988 and McIntosh 2025 also span
  three. The claim was corrected rather than the threshold loosened: spanning
  three strokes is not unique, doing it while holding more than four titles is.
- **G6 — the top two.** The two highest athlete-years in the entire file must
  both belong to the subject, and the best anyone else has managed is reported
  alongside.
- **G7 — the marginal title.** In each record year, the narrowest winning margin
  must be at most 0.05s for the script to call it "a hundredth", and removing
  that single title must drop the count as the script says it does.
- **G8 — "one stroke swum twice".** The script claims two world #1s in a year is
  *almost always* the same stroke repeated. That is a quantitative claim, so it
  is measured; the build fails below 75%.

## Results

Across 1985–2025, both sexes, all 17 individual long-course events, **1,394
event-years** were retrieved, of which **1,184** cleared the coverage floor.
Those produced **865 athlete-years** with at least one world #1.

| world #1s in one year | athlete-years | share |
|---|---|---|
| 1 | 606 | 70.06% |
| 2 | 204 | 23.58% |
| 3 | 44 | 5.09% |
| 4 | 9 | 1.04% |
| 5 | 1 | 0.12% |
| 6 | 1 | 0.12% |

- The **top two athlete-years ever recorded are both Michael Phelps**: 2007 with
  six, 2008 with five. Nobody else has ever exceeded four.
- The nine athlete-years at four are Otto 1988, de Bruijn 2000, Lochte 2011,
  Ledecky 2017 and 2020, Sjöström 2017, Dressel 2019, Marchand 2024 and
  McIntosh 2025.
- **167 of 204** two-title athlete-years (81.9%) are a single stroke family.
- Both of the subject's record years turn on **0.01s**: the 2007 100m freestyle
  (48.42 to Phelps, 48.43 shared by Hayden and Magnini) and the 2008 100m
  butterfly (50.58 to Phelps, 50.59 to Čavić). Remove those two hundredths and
  the record years become five and four — the second of which is merely level
  with the other nine.
- The record survives deleting the 2008–09 supersuit years entirely (G2) and
  survives moving the cutoff from *#1* to *top three* (G4).

## Honest limits

- **The archive is not the sport.** Pre-1995 coverage is thin even for major
  events, and the coverage floor drops those event-years rather than counting a
  shallow one. Anyone ranked #1 in a year the archive barely covers is absent
  from this file, not demoted.
- **Rankings are not head-to-head.** A season world #1 means the fastest time
  swum that year, which is not the same as beating everyone in a final. Most of
  the subject's 2008 marks happen to be Olympic finals, but the metric does not
  require that.
- **Era comparison is deliberately avoided.** Nothing here claims a swimmer from
  one decade would beat one from another. The unit is always "ranked #1 among
  their own contemporaries in that calendar year".
- **Sexes are never pooled** into a single ranked comparison; they are counted
  in the same distribution only as separate athlete-years.
- **The margin is a season margin.** "Won by 0.01s" means the year's fastest and
  second-fastest times differ by 0.01s — in this case both were swum in the same
  Olympic final, which is stated rather than assumed.

## Sources

- World Aquatics official season world rankings, long course metres
  (`api.worldaquatics.com`), retrieved July 2026.
