reelgorithm.py

The worst US airline is not the one that’s late the most

We scored all nine US mainline carriers on 5,232,629 real flight records, with the weights written down before we looked. The airline that finishes last is 7th of 9 at being on time. It finishes last because of the failure that never shows up in the statistic everyone quotes.

Every “worst airline” ranking is a vibe with a logo next to it. That is not a complaint about the taste involved — it is a complaint about the arithmetic. The weights are almost never published, and a ranking whose weights you cannot see is a ranking you cannot disagree with.

So we did it the other way round. Pick the metrics first. Write the weights down. Apply them to all nine carriers identically, from the federal flight records, and let the answer be whatever it turns out to be.

It turned out to be Spirit — and the reason why is much more interesting than the name.

The ruler

Three metrics, weighted equally at a third each. Fixed before the data was pulled, which is the only thing that stops this from being a search for a headline.

the score — three metrics off the BTS microdata, weights locked in advance
metricweightmeasured asbetter
On-time arrivals1/3arrived within 15 min, over flights that operatedhigher
Cancellations1/3cancelled, over flights scheduledlower
Mean arrival delay1/3delay minutes per operated flightlower

Each carrier is scored on its z-score within the nine, not its percentile rank — the opposite of the choice we made when ranking 95 cities, and for a specific reason. With only nine units, a percentile throws away almost everything. It would insist that the gap between 8th and 9th is the same size as the gap between 1st and 2nd. In this field those two gaps differ by an order of magnitude, and that difference turns out to be the story.

The three z-scores are signed so that positive always means better for the passenger, then averaged. A score of −1.09 means “1.09 standard deviations worse than the average US mainline carrier” — not 1.09 out of anything.

One detail does real work: cancellations are measured over every flight scheduled, while on-time is measured only over flights that actually operated. Folding a cancelled flight into the on-time rate as a very late arrival would count the same failure twice. Keeping them separate is what lets the two metrics disagree — and they disagree loudly.

The nine, and why only nine

Southwest, Delta, American, United, Alaska, JetBlue, Frontier, Allegiant and Spirit: every US mainline carrier with twelve complete months of data, June 2025 through May 2026.

The regional operators are excluded, and it is worth knowing why. BTS reports the operating carrier, so SkyWest, Envoy, Republic and PSA appear under their own names rather than as “United Express” or “Delta Connection”. You cannot book a SkyWest ticket, and ranking a company most passengers have never heard of alongside the brands would answer a question nobody asked. The honest cost of that decision: a mainline carrier’s regional flying is not charged to it here. Hawaiian is out too — it reports only 7 of the 12 months after the Alaska merger, and a partial year silently changes every rate.

The answer

all nine, best to worst — 5,232,629 flights, June 2025 – May 2026
#carrierscoreon-timecancelledmean delay
1Alaska+0.9478.25%1.30%12.5 min
2Delta+0.9480.93%1.67%14.6 min
3Southwest+0.8977.13%0.89%13.7 min
4United+0.8678.98%1.01%16.2 min
5Allegiant−0.0874.20%0.63%23.1 min
6American−0.7074.27%2.28%22.8 min
7Frontier−0.8672.52%1.93%24.0 min
8JetBlue−0.9072.53%2.50%21.7 min
9Spirit−1.0973.40%3.41%20.8 min

Spirit is last. It is also 7th of 9 on being on time — ahead of both Frontier and JetBlue. Its mean delay is 5th, comfortably mid-table. On two of the three metrics it is unremarkable.

It finishes last on the strength of one number. Spirit cancelled 3.41% of everything it scheduled, against a nine-carrier average of 1.74% and a best-in-field 0.63%. That is 2.0 standard deviations below the mean — the largest single-metric deviation anywhere in the table, in either direction. Nothing else in this dataset is that far from normal.

The margin over 8th-placed JetBlue is 0.225σ. That is narrow, and it deserves saying plainly rather than dressing up: Spirit is the worst airline in America by this ruler, but not by much.

The turn — punctuality is not the ranking

Here is the part that makes the whole exercise worth doing. The metric every airline advertises, and every ranking leans on, does not agree with the full score:

where the two rankings cross
carrieron-time rankoverall rank
Frontier9th — worst in America7th
Spirit7th9th — worst overall
American5th6th
Allegiant6th5th

Frontier has the worst on-time rate in the United States and is not the worst airline. Spirit is on time more often than Frontier and still finishes last.

The mechanism is simple once you see it, and slightly infuriating afterwards: a cancelled flight never gets to be late. It does not arrive 90 minutes behind schedule and drag the average down — it leaves the on-time statistic altogether, because there is no arrival to measure. An airline having a catastrophic day can therefore improve its headline punctuality number by cancelling the flights that were going to be disasters.

We are not claiming anyone does that deliberately. The point is structural: judge airlines on the number they all quote, and you systematically under-punish whichever ones resolve their worst days by cancelling rather than by flying late. The passenger whose flight vanished is not in the statistic at all.

The shape of the field

The nine do not spread out evenly. They clump, and the clumping is more informative than most of the individual positions.

three groups, and one very large hole
groupwhointernal spread
The top fourAlaska, Delta, Southwest, United0.086
— the gap —United down to Allegiant0.936
Alone in the middleAllegiant
The bottom fourAmerican, Frontier, JetBlue, Spirit0.395

The hole between 4th and 5th is roughly eleven times the entire spread of the top four. Alaska, Delta, Southwest and United are, for practical purposes, the same airline as far as this ruler can tell — Alaska beats Delta by 0.001, which is nothing. Arguing about which of the big four is best is arguing about noise. The distance from that group to everybody else is the real finding.

Allegiant, sitting by itself, is the most interesting carrier in the table: best in the entire field on cancellations at 0.63%, and second-worst on delay at 23.1 minutes. It almost refuses to cancel, and is late a great deal instead. That is the exact opposite trade to Spirit’s, and it is why the two sit where they do.

Does the answer survive?

The delay metric has a defensible alternative definition. We measured mean delay per operated flight, where an on-time arrival counts as zero. You could instead measure the average delay of a delayed flight, which is a much bigger number — Alaska is 12.5 minutes on our definition and 57.6 minutes on that one.

Since the choice is arguable, we recomputed the entire ranking under the alternative:

the robustness check — same ruler, other delay definition
delay defined as…worst airlinemargin over 8th
per operated flight (what we shipped)Spirit0.225σ
conditional on being lateSpirit0.265σ

Spirit is last either way, and the margin actually widens slightly under the alternative. The headline does not depend on the judgement call.

The middle of the table does. Frontier and JetBlue swap, and so do Delta and United. Which is one more reason to treat any individual mid-table position as noise, and a reason we are publishing both the ruler file and all 108 carrier×month rows rather than just the answer.

What happened to Spirit

The twelve months in this dataset are not an ordinary year for the company, and the monthly file shows it:

Spirit, scheduled flights per month
June 2025May 2026
Flights scheduled17,623354

A 98.0% reduction in scheduled flying across the window. The worst single month for punctuality was March 2026, at 54.31% on time. The worst for cancellations was May 2026 at 14.12% — roughly one scheduled flight in seven simply not operating.

One caution, because it would be easy to over-read: 354 flights is a very small denominator, so that final month’s rates are volatile and should not be treated as a stable estimate of anything. The collapse in volume is the robust fact. The rate in the last month is not. The twelve-month composite is dominated by the months when Spirit was still flying at scale, which is the correct behaviour — it is a flight-weighted picture rather than a month-weighted one.

What is deliberately not in here

the honest limits
left outwhy
PriceThis measures whether the airline did the thing it sold you, not whether it was worth the fare. A score with price in it would mostly be measuring “is this a budget carrier”, which you already know
Legroom, service, bagsNot in the flight records, and every public index of them is somebody’s opinion weighted by somebody else’s. Three measured metrics beat six with three invented
SafetyNot measurable at this resolution. Fatal accidents on US mainline carriers over twelve months are so rare that any rate built from them is noise. Leaving it out is not a claim that all nine are equally safe — it is a refusal to pretend a one-year window can answer that
Regional flyingCharged to the regional operator, not to the brand that sold the ticket. A real limitation of using operating-carrier data

The weights are a judgement. They are published so that you can disagree with them, and the files below carry every raw count and every z-score, so disagreeing is a spreadsheet exercise rather than an argument. If your weights put Spirit third, that is a real answer too — and you will be able to show your working, which is more than the lists can.

the data — everything above is reproducible from these

us-airline-ruler-2026.csv — all nine carriers, the composite score, the overall rank, each metric’s measured value and rank, all three z-scores, and the raw scheduled / operated / cancelled counts behind every rate.
us-airline-monthly-2026.csv — 108 rows, one per carrier per month. This is the working data the ranking is built from, shipped so you can check the answer or find something we missed.
us-airline-ruler-methodology.md — the keyless BTS pull, the exact definition of every metric, the robustness check in full, and what was left out and why.