reelgorithm.py

Your second most-pressed key is not a letter

We measured finger travel across 65.8 million characters of permissively licensed source code and 5.7 million of public-domain prose, on a standard QWERTY board. Code moves your fingers 1.267× farther per 1,000 characters. Your right pinky goes from an eighth of the work to nearly a third. And the second most-pressed key in all of code is Left Shift.

QWERTY was arranged for English. Not for a good reason — the layout is a settlement between mechanics, telegraphy and habit — but it was arranged, and the thing it was arranged around was prose. Then a few million people started using it for eight hours a day to type something else.

So: how much does that actually cost, in millimetres?

What we measured

One key unit is 19.05 mm, the standard ANSI key pitch. Put the keys on that grid, assign fingers by touch-typing convention, and walk a text character by character. Each finger travels from wherever it last landed to the key it has to press next.

That last clause is the only interesting modelling decision here. Fingers do not snap back to home between keystrokes. A return-to-home model roughly doubles every distance, and it also changes the answer, because it erases the cost of a finger being dragged off its home position and kept there. Which is most of what punctuation does.

A capital or a shifted symbol counts as two presses: the character, and a Shift pressed by the opposite hand’s pinky. Hold on to that — it is the whole article.

Prose costs 14,686 mm. Code costs 18,612.

Per 1,000 characters. That is the 1.267×, and it comes from ten Project Gutenberg novels against 65 repositories — thirteen languages, five repos each, all MIT, Apache, BSD or public domain.

The gap is not spread evenly across your hands. It lands almost entirely on one finger:

Your right pinky does 12.4% of all finger travel in prose and 30.8% in code. It owns the bracket, the brace, the quote, the colon, the semicolon and the hyphen. In English those are garnish. In code they are the grammar.

And then there is Left Shift

Rank the keys of a standard board by share of all presses in our code corpus and the top of the list is: space at 19.97%, Left Shift at 8.35%, then e at 6.61%.

The second most-pressed key in source code is not a letter. It beats E, which is the most common letter in English by a wide margin, and only the space bar is ahead of it.

The mechanism is the two-press rule. You press Left Shift with your left pinky for characters your right hand is typing — and code capitalises and punctuates relentlessly. Every {, every :, every ", every _, every CamelCase boundary, every CONSTANT. Shift presses go from 3.0% of all presses in prose to 11.9% in code.

The part where the story does not work

The obvious next move is a leaderboard. Thirteen languages, rank them, crown the worst one. We built it, and then we had to throw it away.

Language does matter in aggregate — η² = 0.598, with a permutation p of 0.00005. But the ordering is mostly noise. Only 1 of 12 adjacent pairs in the ranking is statistically separable, and 35 of 78 pairs overall. The average spread within a language, across its five repos, is 2,684 mm — against a total range between languages of 5,740 mm.

Put plainly: swap one Python repo for another and you move further than the distance between most neighbouring languages. A 13-bar chart would have looked authoritative and been an artefact of which repos we happened to pick.

This is also why five repos per language and not one. An earlier version of this scored a single repository per language, which does not measure a language at all — it measures one team’s brace convention, comment density and identifier length, wearing a language’s name.

Exactly one language survives

Go, at 22,245 mm per 1,000 characters. It is separable from all twelve others, and all five Go repos land inside 21,726–22,865 mm, which is a tight band by the standards of this dataset.

Testing that honestly takes some care. A Bonferroni correction is unreachable here by construction: 0.05/78 = 0.00064, but an exact five-versus-five permutation test cannot return a p below 2/252 = 0.0079. No pairwise correction can ever pass, so a pairwise claim would be theatre. Instead we tested the headline jointly — does any language separate from all twelve others?, under 2,000 relabellings — which gives p = 0.0010.

So Go, against a band. Everything below it is a tie wearing a ranking.

The part that argues against us

The big one, and it is big: this is per 1,000 characters, not per unit of work. Code says more per character than prose does. So this is the cost of typing code, not of writing it, and a terser language could cost more per character while costing less per program. If you want to overturn the framing, start here.

Two claims we had to kill on the way, both worth more than the finding they replaced. “Even the cheapest code source beats the worst prose” is false — fastapi at 13,494 mm is cheaper than Alice in Wonderland at 15,208 mm. And “Go’s capitalised identifiers drive its shift load” is falsified: Go is third in Left-Shift presses, behind C++ and C.

One more, because the error generalises. Replacing every tab with four spaces in the Go repos saves 6.6% of absolute travel. Our first pass said 18%, because it normalised per 1,000 characters — and the four spaces the intervention adds also inflate the denominator, so the saving gets counted twice. Normalise a counterfactual on the quantity the intervention does not change.

Finally: no human typed anything. This is a geometric model over text. There is no timing, no error rate, and nothing here is an ergonomic or medical claim.

Get the data

All 75 sources — every novel and every repository with its licence, character count, travel, shift share and pinky share — plus the 13-language table and a per-key breakdown of all 53 keys. Free, no email required.

Open the data kit

Sources: ten Project Gutenberg texts (US public domain, boilerplate stripped) and 65 repositories under MIT, Apache-2.0, BSD, BSL or Unlicense, each recorded with its licence in the data file. Nothing upstream is redistributed — the download is a table of counts and distances we computed. Keyboard model: ANSI QWERTY, 19.05 mm key pitch, no return-to-home. Analysis and figures: reelgorithm.py.