# Which US city is actually the best place to live? — dataset & methodology

*95 US metros · 4 federal datasets · one composite score, with the weights
written down before anything was ranked · and a fifth metric deliberately left
out because it could not be measured.*

**`us-city-ruler-2026.csv`** — 95 rows, one per metro, sorted best to worst.

| column | meaning |
|---|---|
| `rank` | 1 = best. Derived from the score, never re-sorted separately |
| `metro`, `cbsa`, `cbsa_name` | the metro, its Census CBSA code, and the official title |
| `score` | the composite, 0–100 |
| `cost_value`, `cost_rank` | BEA Regional Price Parity, all items, 2023. Index, US = 100. **Lower is better** |
| `housing_value`, `housing_rank` | % of ALL households paying 30%+ of income on housing, ACS 2023 5-yr. **Lower is better** |
| `jobs_value`, `jobs_rank` | Total nonfarm employment, latest 12-month % change, BLS. **Higher is better** |
| `commute_value`, `commute_rank` | mean one-way commute, minutes, ACS 2023 5-yr. **Lower is better** |
| `net_domestic_migration_per_1k_2024` | Census PEP. **Not part of the score** — see "the turn" |
| `migration_rank` | 1 = most people moving in, over the same 95 |
| `population_2024` | Census PEP |

Every `*_rank` is a **midrank**: tied metros share the average of the ranks they
span, so `2.5` means two metros tied for 2nd. This is not cosmetic — see
"the tie bug" below.

---

## The question

Every "best places to live" list is somebody's opinion with a number bolted on.
The weights are usually invisible, which is what makes them unfalsifiable. So:
pick the metrics first, publish the weights, apply them to every metro
identically, and let the answer be whatever it is.

## The ruler

Four metrics. The weights were fixed before the data was pulled.

| metric | weight | source | direction |
|---|---|---|---|
| Cost of living | 30% | BEA Regional Price Parities, all items, 2023 (US = 100) | lower better |
| Housing burden | 25% | Census ACS 2023 5-yr, B25091 + B25070 | lower better |
| Job growth | 25% | BLS Total Nonfarm, latest 12-month % change | higher better |
| Commute | 20% | Census ACS 2023 5-yr, B08303, mean one-way minutes | lower better |

**Normalisation is percentile rank within the 95, not z-score.** The raw units
are not comparable to each other, and two of them are badly skewed — San
Francisco's price level and New York's commute would drag everyone else's
z-score toward the middle and compress the field. A percentile is immune to
that. The cost is that it throws away the *size* of a gap: being 1st on housing
scores the same whether you beat 2nd place by a hair or by ten points.

The composite is the weighted sum of the four percentiles, ×100.

**Housing burden is owners AND renters.** B25091 (owners with a mortgage) alone
would drop about 35% of households, and renters are the more cost-burdened half,
so owners-only is not the metric it looks like. B25091 + B25070 together are the
standard "cost-burdened" definition.

## What is NOT in it

**Violent crime.** It belongs in this score and it is not here. The FBI's Crime
Data Explorer API works without a paid key, but exposes **state and agency
endpoints only — there is no metro endpoint**. Building a metro crime rate means
aggregating thousands of individual police agencies through a county→CBSA
crosswalk, which the demo key's rate limit makes impossible. It needs a free
api.data.gov key and a day of work.

So the ruler is four metrics and says so on screen. **Four measured metrics beat
five with one invented.** No placeholder was substituted, and none should be.

Also absent, and deliberately: schools, weather, culture, healthcare. Each is
either unmeasured at metro level, or measurable only as somebody's index of
somebody else's index.

## The 95 metros

The 95 largest US metros by 2024 population, floor ~464,000 (Springfield MA),
totalling **223.5 million people — about two thirds of the country**. The
selection is by size alone. It is not a shortlist of nice places, which matters:
a ranking of places somebody already chose is not a ranking.

Metros, not cities. "Austin" here is the Austin-Round Rock-San Marcos CBSA,
because commute times and price levels are properties of a labour market, not of
a city limit.

## How the data was pulled

    BEA RPP.zip                    ->  price level, 2023          [no key]
    BLS SM series, download.bls.gov ->  12-month job growth       [no key]
    Census ACS table-based SF       ->  B25091 + B25070, B08303   [no key]
    Census PEP cbsa-est2024         ->  net domestic migration    [no key]

**None of these need an API key.** `api.census.gov` now refuses keyless requests
(302 → `missing_key.html`), and it is easy to conclude from that the data is
gated. It isn't — every one of these agencies publishes the same numbers as a
flat file that needs nothing. That detour cost a day; it is written down here so
it costs you none.

Metros are matched to CBSA codes off the **official Census geography file**, not
a hand-written table (see "matching the wrong real city" below). Two metros were
bridged across a redelineation between the BEA (2020) and ACS (2023) vintages:
Poughkeepsie → Kiryas Joel-Poughkeepsie-Newburgh, and Cleveland-Elyria →
Cleveland.

## The finding

**Fayetteville-Springdale-Rogers, Arkansas — 96.4 / 100.** 4.6 points clear of
#2 Wichita, and 2.25 standard deviations above the 95-metro mean of 50.0
(sd 20.7).

| metric | measured | rank | points earned |
|---|---|---|---|
| Cost of living | 91.0 (US = 100) | 8th | 27.8 / 30 |
| Housing burden | 23.1% | **1st** | 25.0 / 25 |
| Job growth | +2.87% | **2nd** | 24.7 / 25 |
| Commute | 22.8 min | 6th | 18.9 / 20 |

**The mechanism is the absence of a weakness, not the presence of a strength.**
Fayetteville never falls below **8th of 95 on any of the four**, and no other
metro in the set manages that. A weighted sum of capped percentiles punishes a
hole harder than it rewards a win — you can only earn 30 points on cost no
matter how cheap you are, but you can lose all 30. So the winner is not the best
at anything in particular. It is the only one with nothing wrong with it.

The runners-up have the same shape: Wichita never below 18th, Omaha never below
27th.

**The field is tight.** From 10th place to 20th is **7.0 points** of score. That
is the real reason every published "best places" list disagrees with the next
one — in the middle of the distribution, a trivial change of weights reshuffles
twenty cities. The distinction between #12 and #19 is not a finding. The
distance from the pack to #1 is.

Full range: 13.1 (Miami) to 96.4.

## The turn — the score against where people actually move

A score nobody can disagree with proves nothing, so it is tested against
something measured and **deliberately excluded from the score**: Census net
domestic migration, 2024, per 1,000 residents, over the same 95 metros.

**Spearman ρ = 0.50.** Real agreement, and a long way from the same thing. The
divergences are the interesting part:

| metro | ruler rank | migration rank |
|---|---|---|
| Lakeland, FL | 55th | **1st** |
| North Port, FL | 61st | 2nd |
| Jacksonville, FL | 67th | 6th |
| Salt Lake City, UT | 16th | 85th |
| Toledo, OH | 19th | 73rd |
| El Paso, TX | 21st | 89th |

Both ranks are computed over the **same 95 metros**. Ranking a subset against a
full set and calling the difference a finding is the easiest mistake in this
genre and it is worth naming.

This is not the ruler being wrong. It measures cost, housing, jobs and commute —
it does not measure weather, family, or the fact that a place is where you are
already from. Migration is people optimising for a different function.

## Three failures that produced correct-looking charts

Kept because each one is invisible in the output.

**1 · The tie bug.** BEA publishes the price index to one decimal, so ties are
common. Sorting to get a percentile gave two metros with *byte-identical
measured values* scores up to **1.5 points apart**, depending only on where they
sat in the file. A ranking that depends on file order is not a ranking. Fixed
with the midrank, `pct = (worse + (tied − 1) / 2) / (n − 1)`, and the displayed
rank is inverted back out of the percentile rather than re-sorted, so the rank
shown can never disagree with the score that produced it. This changed the
winner's score from 96.8 to 96.4 and its worst per-metric rank from 7th to 8th.

**It was caught only by writing the scoring a second time, independently, and
diffing all 95 scores.** Nothing else would have found it. The two
implementations now agree to 0.

**2 · A sign error in the normaliser** ranked Miami #1 on a cost index of 111.8
(89th of 95) and a housing burden of 45.2% (94th of 95). The chart looked
completely normal — bars in a sensible order, a plausible winner. Direction is
now declared once, in one table, and used in exactly one place.

**3 · Matching the wrong real city.** Indexing CBSAs in file order matched "Las
Vegas" to Las Vegas, **New Mexico**, and "Cleveland" to Cleveland, **Tennessee**
— micro areas that share a principal city name with a major metro and appear
earlier in the file. Nothing errors. You get real, correct, verifiable numbers
for the wrong city. Fixed by matching off the official Census geography file and
printing the full mapping for a human to read.

## What would change the answer

- **Adding crime.** It is the biggest known gap. Fayetteville's margin is large
  enough that it would probably survive, but "probably" is not a result.
- **Weighting differently.** The weights are a judgement, published so you can
  disagree with them. The CSV has every raw value and every per-metric rank in
  it, so you can re-weight the whole thing yourself in a spreadsheet.
- **Wages.** Cost of living without income is half a picture. A cheap metro with
  no jobs scores well on 30% of this ruler. Job *growth* partly covers it;
  median earnings would cover it properly.

## Sources

- BEA Regional Price Parities, 2023 — `apps.bea.gov/regional/zip/RPP.zip`
- BLS State & Metro Area Employment (SM), Total Nonfarm — `download.bls.gov/pub/time.series/`
- Census ACS 2023 5-year, B25091 / B25070 / B08303 — `www2.census.gov/programs-surveys/acs/summary_file/2023/`
- Census Population Estimates, CBSA totals 2020–2024 — `www2.census.gov/programs-surveys/popest/datasets/2020-2024/metro/totals/`

Data generated 2026-08-14. Everything above is reproducible from the CSV.
