# DSWAGE_METHODOLOGY.md — the data-science wage premium, by US metro

The number behind the episode. Built 2026-09-17 from
`scripts/ds_wage_fetch.mjs` + `scripts/ds_wage.py`. The artifact is
`data/dswage/ds_wage_premium.csv`; the episode's copy reads
`public/dswage/ds_wage.json` and **may not say a number that is not in it**
([[feedback_never_fabricate_data]], [[feedback_numbers_come_from_the_artifact]]).

Supersedes the AXIS subject for this slot — `BIOCLICK_HANDOFF.md` §0 was the
blocking decision and the user resolved it 2026-09-17 in favour of the pivot.

---

## 1 · The question

> Where in America is being a data scientist actually worth the most?

Not "which metro pays the highest salary" — that is a Googleable table and it
answers a different question. A $173,160 salary in San Jose is earned in a metro
where the **typical job already pays $82,470**. The same title in Charlotte pays
$131,110 against a local median of $48,880.

So the measure is the **local wage premium**:

```
premium(metro) = median annual wage, data scientists  ÷  median annual wage, all jobs
                 (same metro, same year, same survey)
```

⚠️ **THIS IS WHY NO COST-OF-LIVING INDEX IS NEEDED.** A ratio of two wages
inside one metro is already denominated in local prices — the price level
divides out algebraically. That is the methodological point of the episode, and
it is what lets the analysis sidestep BEA RPP's CBSA vintage bridge entirely.
RPP is still computed (`real_median` in the CSV) but **only** to show the naive
cost-of-living cut that beat 2 corrects. It is not what anything is ranked on.

## 2 · The data

| file | what | bytes, verified 2026-09-17 |
|---|---|---|
| `oesm24ma.zip` | BLS OES metro, May 2024 | 40,189,821 |
| `oesm23ma.zip` | BLS OES metro, May 2023 | 39,224,418 |
| `RPP.zip` | BEA Regional Price Parities, 2023 | 252,119 |
| `oesm24st.zip` | BLS OES **state** file, May 2024 | 7,617,815 |
| `gaz_cbsa.zip` | Census 2024 Gazetteer, CBSA centroids | 46,177 |

OES gives every SOC occupation × every CBSA: employment, location quotient, and
the mean/median/10th/25th/75th/90th wage. Occupation is **15-2051 Data
Scientists**. The all-jobs denominator is the same file's `O_GROUP == 'total'`
row for that metro, so numerator and denominator come from one survey and one
year and cannot drift apart.

393 metros are in the file.

## 3 · The pre-registered inclusion rule

Decided **before** the winner was looked at, and stated on screen because it
changes the answer:

1. `TOT_EMP >= 100` data scientists in the metro.
2. A published median wage in **both** May 2023 and May 2024.

**135 metros** survive. 17 are dropped by rule 2 and are written to
`data/dswage/excluded_one_year.csv` rather than silently disappearing.

⚠️ **RULE 2 EXISTS BECAUSE THE TOP OF A NOISY RANKING IS A WINNER'S CURSE.**
Ranking ~150 survey estimates by their maximum selects whoever drew luckiest,
not whoever is highest. The single-year 2024 table is topped by **Idaho Falls,
ID at 3.64×** — a real and tightly-estimated number (220 data scientists, mean
PRSE 1.8%, Idaho National Laboratory is the mechanism) — but its 2023 wage is
suppressed (`*`), so there is exactly one year of it and nothing to check it
against. It is reported, not hidden, and it is not the answer.

## 4 · The three traps in the source

Recorded in `mcp/src/sources.json → bls-oes-metro`; all three are silent.

1. **`www.bls.gov` refuses `curl`** — 403 bare *and* with a browser User-Agent.
   Node's `fetch` gets through with no headers at all. ⊥ reach for a spoofed UA:
   on `cpsc.gov` a spoofed UA is what *causes* the block. Hence
   `ds_wage_fetch.mjs` is Node, not Python `requests`.
2. **`#` means "at or above $115.00/hr", NOT missing.** Treating it as null
   silently deletes the highest-paid metros and inverts the answer. In May 2024
   the `#` rows in `A_PCT90` are exactly **San Jose and Boulder** — i.e. a
   null-drop erases the top of the country. `*` is the genuinely-suppressed
   marker and *is* null. The script encodes `#` as the floor 115.00 × 2080 =
   $239,200 and flags it as `p90_topcoded`.
3. **BEA RPP is on 2020 CBSA delineations, OES on 2023.** Two of the 95 largest
   metros moved: `39100 → 28880` (Poughkeepsie → Kiryas Joel) and
   `17460 → 17410` (Cleveland-Elyria → Cleveland). Bridged explicitly.

## 5 · The answer

```
WINNER   Charlotte-Concord-Gastonia, NC-SC
         $131,110  vs local median $48,880  =  2.68×   (2023: 2.79×)
         3,870 data scientists · mean PRSE 1.5% · LQ 1.91 · raw wage rank 5

NAIVE    San Jose-Sunnyvale-Santa Clara, CA
LEADER   $173,160  vs local median $82,470  =  2.10×
         raw wage rank 1  →  premium rank 50 of 135
```

Median premium across the 135 metros: **2.02×**. Year-over-year rank stability:
**Spearman ρ = 0.750**.

Charlotte is first on the two-year mean (2.73×) *and* first in each year taken
alone (2.785× in 2023, 2.682× in 2024), on a 3,870-person cell — so it is not a
small-sample artifact. The mechanism is that Charlotte is a banking centre with
a low general wage base: the numerator is a finance-sector data salary and the
denominator is a normal Southern metro.

⚠️ **THE EPISODE'S TURN IS SAN JOSE'S 49-PLACE FALL** (rank 1 → 50 of 135),
which is the same shape as NFLRobbed's: the hook spikes on the naive measure and
the analysis takes it apart.

## 5b · The green choropleth under the hook

The hook's states are filled green in proportion to what a data scientist makes
in that STATE, brighter = more (user, 2026-09-17). Range **$69,430 (Mississippi)
→ $158,760 (Washington)** across 51 states.

⚠️ **FROM THE OES STATE FILE, NOT AGGREGATED FROM THE METRO TABLE.** Rolling the
135 ranked metros up to states would drop every non-metropolitan area and tilt
each state toward its cities — a different quantity wearing the same name.

⚠️ **DELAWARE IS NEUTRAL GREY, NOT DARK.** It has 540 data scientists and a
suppressed median (`*`). Painting it at the bottom of the ramp would say, in the
encoding a viewer reads fastest, that Delaware pays the least in America. The
legend carries a "no data" swatch for it.

⚠️ **TWO QUANTITIES AT TWO GEOGRAPHIES ARE ON SCREEN AT ONCE** — spike height is
a METRO median, the fill is a STATE median — so the hook labels both. An
unlabelled second encoding is a chart lying by omission.

⚠️ **GREEN IS A DEPARTURE FROM THE HOUSE PALETTE AND WAS ASKED FOR.** This
channel is pink-on-near-black with no third hue ([[feedback_codepop_style]]).
Pink stays reserved for the thing being MEASURED, so the green is ground and
never an accent, and no other beat uses it.

⚠️ **THE RAMP RUNS `#03120A` → `#6DFF9E`, AND IT IS ONLY THAT WIDE BECAUSE THE
SPIKES CAME OFF** (2026-09-18). It was capped at a muted `#3FA96A` for exactly
one reason — white pins stood ON the fill and their lower halves vanished into a
bright one, the same trap NFLRobbed's `FILL_MAX` guards against. With nothing
drawn on top, the full range is available: about a 15× luminance span against
the old 4×.

⚠️ **THE INTERPOLATION IS LINEAR IN OKLab AND MUST STAY THAT WAY.** `mixOk` is
perceptually uniform, so a linear t gives a perceptually linear ramp — equal wage
steps look like equal colour steps. Adding a gamma to "boost contrast" would make
the picture overstate differences at one end of the range, a lie factor above 1
on the frame most people will ever see ([[feedback_phd_chart_standard]]). **Widen
the endpoints; never bend the middle.**

## 6 · The kill gates — `IDEA_GENERATION.md`

**KILL 1 — the banger frame.** ✅ Frame 0 is the tilted US map already pulling
back with a spike on every one of the 135 metros, height = raw median wage, San
Jose towering over everything. That is the shot ported from
`series/nflrobbed/Beat1Map.tsx`, which is the only form measured at low skip@3s
(NFLFan 33.6% / 443k). It is in motion at t=0 and legible with no setup.

**KILL 2 — never-been-made.** ✅ The *subject* (data-science salaries by city) is
proven and heavily covered; the amended gate allows that. The *analysis* is
ours: nobody has published the OES per-metro premium against the same metro's
all-jobs median, under a two-year pre-registration rule, with the `#` top-code
handled. The un-Googleable question is "where is this job worth the most
**relative to your neighbours**", not "where does it pay most".

**Hard rules.** 6 US-monetizable ✅ (US metros, US dollars, wages — explicitly a
high-CPM subject). 8 zoomed-out visual ✅ (135 metros on a national map). 9
intellectually interesting ✅ (a ratio that cancels the price level, a
pre-registered inclusion rule, a winner's-curse correction, a top-code trap).
10 real insight ✅ (nobody guesses Charlotte, and "the Bay Area premium is
ordinary" is the opposite of the received view).

**Value gate.** V1 forwardable 5 — *"A data scientist in Charlotte earns 2.7×
the typical local wage. In San Jose it's 2.1× — 50th in the country."*
V2 valuable 4 (nominal vs relative wage is a real mental model).
V3 personal stakes 5 (money, career, where you live).
`KILL1 ✅ KILL2 ✅ | V1 5 / V2 4 / V3 5 = 14 ✅`

**Surprise type:** *the ranking inverts under the right denominator* — a
controlled comparison, not "the average is a lie".

## 7 · Caveats that must survive to the screen

- OES is May 2024, published 2026; RPP is 2023. Where RPP appears it is a year
  behind the wages and is labelled as such.
- The premium is a ratio of **medians**, not of the same people's pay. It says
  what the job is worth against the local wage base, not what one person's
  raise would be for moving.
- Metro-level `TOT_EMP` for a single SOC code is an estimate with its own error;
  `prse` is carried in the CSV for every row.
- 17 metros are not ranked. Idaho Falls would lead on one year of data.

## 8 · Reproduce

```bash
node scripts/ds_wage_fetch.mjs                      # fills .cache/dswage/
../.venv/bin/python scripts/ds_wage.py              # writes data/ + public/
```
