# QWERTY finger travel — methodology

The measurement behind *"Which language makes your fingers travel farthest?"*
Written 2026-09-21.

**The code is the specification.** `scripts/qwerty_load.py` and
`scripts/qwerty_layout.py` are the only places anything is computed; this
document explains them rather than restating them. Every random draw is seeded
(`SEED = 20260920`), so a rerun reproduces these numbers exactly.

**Nothing here redistributes a source.** These files are a table of counts and
distances *we computed*, so the download carries no upstream terms at all.

---

## 1 · What is being claimed

**Typing source code moves your fingers 1.267× farther than typing English
prose does, per 1,000 characters.** Prose costs **14,686 mm** per 1,000
characters; code costs **18,612 mm**.

What it is **not**: a claim that code is harder to write, or that any language
is bad. It is a distance, under one keyboard model, per *character* — see the
caveat in §6, which is the most important section here.

---

## 2 · The keyboard model

A standard **ANSI QWERTY** board. One key unit `1u = 19.05 mm`, the standard
key pitch. Keys sit on a grid: number row `y=1`, top `y=2`, home `y=3`,
bottom `y=4`, space `y=5`, with the usual per-row horizontal stagger (top row
offset 1.5u at Tab, home 1.75u at Caps, bottom 2.25u at Left Shift).

Fingers are assigned by touch-typing convention — eight fingers on the home
row, thumb on space.

⚠️ **Fingers do not snap back to home between keystrokes.** Each finger travels
from wherever it last landed to the key it must press next. This is the single
biggest modelling choice in the project. A return-to-home model roughly doubles
every distance and, more importantly, *changes the ranking*, because it erases
the cost of a finger being dragged away from its home position and kept there.

A capital letter or a shifted symbol is **two presses**: the character, and a
Shift pressed by the **opposite hand's pinky**. That convention is what makes
Left Shift expensive, and it is why Left Shift — not a letter — is the second
most-pressed key in code.

## 3 · The corpora

| | |
|---|---|
| **prose** | 10 Project Gutenberg books, US public domain, PG boilerplate stripped — 5,664,121 characters |
| **code** | 65 repositories, **13 languages × 5 repos**, all permissive (MIT / Apache-2.0 / BSD / BSL / Unlicense) — 65,758,519 characters |

Every source, with its licence and its own measured numbers, is in
`qwerty-finger-travel-by-source.csv`.

⚠️ **Five repositories per language, not one.** Until 2026-09-20 this scored a
single repo per language — which measures that repo's house style (brace
convention, comment density, identifier length) and not the language. A
13-way ranking built on n=1 is thirteen engineering teams wearing language
names. Repos were chosen by, in order: permissive licence; **five different
owners**, so no language leans on one org's style guide; different domains;
and wide use.

Three caps keep one project from owning its language: **1,200,000 characters
per repo**, **500,000 bytes per file** (bigger is a generated table), and a
**400-character mean line length** (wider is minified).

Coverage is 99.996% of prose characters and 99.957% of code characters; the
remainder is non-ASCII that a US keyboard cannot type with one key.

**Weighting:** prose is equal per book. The code headline is equal **per
language** — each language is the mean of its five repos, and the 13 language
means are then averaged — so a language with wordier repos cannot buy influence
by being wordier.

## 4 · The headline numbers

| | prose | code |
|---|---|---|
| travel per 1,000 chars | 14,686 mm | **18,612 mm** |
| key presses per char | 1.032 | 1.137 |
| home-row presses | 19.7% | 14.5% |
| Shift presses | 3.0% | **11.9%** |
| **right pinky share of travel** | **12.4%** | **30.8%** |

The right pinky is the story. It owns the bracket, the brace, the quote, the
colon, the semicolon and the hyphen, and it goes from an eighth of the work to
nearly a third.

**Top keys in code by share of all presses:** space 19.97%, **Left Shift
8.35%**, `e` 6.61%, `t` 5.13%. Per-key figures for all 53 keys are in
`qwerty-key-load.csv`.

## 5 · ⚠️ There is no 13-way ranking, and that is the finding

Language matters **globally** — η² = 0.598, permutation p = 0.00005. But the
ranking it implies is mostly noise:

- **only 1 of 12 adjacent pairs separates**, and 35 of 78 pairs overall
- mean within-language spread **2,684 mm** against a between-language range of
  **5,740 mm** — a ratio of 0.468

**Exactly one language can be named: Go**, at **22,245 mm** per 1,000
characters, separable from all twelve others. All five Go repos land in
21,726–22,865 mm. The joint, multiplicity-aware test — *does any language
separate from all twelve others?*, under 2,000 relabellings — gives
**p = 0.0010**.

⚠️ **Bonferroni is unreachable here by construction**, and this is worth
internalising before you re-run it: 0.05/78 = 0.00064, but an exact
five-vs-five permutation test has a p-floor of 2/252 = 0.0079. No pairwise
correction can ever pass. That is why the headline is tested *jointly* instead.

So: Go against a band, never 13 ranked bars. Everything below Go is a tie
wearing a ranking.

## 6 · The caveats that argue against all of it

1. **This is per 1,000 CHARACTERS, not per unit of work.** Code says more per
   character than prose does. So this is the cost of *typing* it, not of
   *writing* it, and a language that is terser may cost more per character
   while costing less per program. This is the caveat most likely to overturn
   the framing, and it is stated on screen.
2. **"The cheapest code beats the worst prose" is FALSE.** It was in an earlier
   draft. fastapi at 13,494 mm is cheaper than *Alice in Wonderland* at
   15,208 mm. The corpora overlap; only the aggregates separate.
3. **Go's shift load is not what drives its lead.** The obvious story —
   capitalised identifiers mean exported symbols — is falsified: Go is **3rd**
   in Left-Shift presses, behind C++ and C.
4. **No human typed anything.** This is a geometric model over text, not a
   study of typists. There is no timing, no error, no ergonomic claim, and
   nothing here says anything about injury.
5. **Repo selection is a judgment call.** The rules in §3 were fixed before the
   answers were looked at, but five repos is five repos. Every per-source row
   ships so you can drop ours, substitute your own, and re-run.

## 7 · The tab counterfactual, and a trap in it

Replacing every tab with four spaces in the five Go repos changes **absolute**
travel from 52,508 m to 49,038 m — **−6.6%**.

⚠️ An earlier pass reported **−18%**, from a per-1,000-character framing. That
is wrong, and the error generalises: the four spaces the intervention
substitutes also **inflate the denominator**, so the saving is counted twice —
once in the numerator, again in a character count padded with three nearly-free
keystrokes. **Normalise a counterfactual on the quantity the intervention does
not change.**

## 8 · The files

| file | what |
|---|---|
| `qwerty-finger-travel-by-source.csv` | 75 rows — every book and every repo, with licence, stars, characters, travel, shift share, pinky share |
| `qwerty-finger-travel-by-language.csv` | 13 rows — per-language mean, SD, min, max and spread |
| `qwerty-key-load.csv` | 53 rows — per key: position, finger, hand, press share and travel in both corpora |

Columns are self-describing. `travel_mm_per_1k` is the measure throughout.
