The shape of a face
There is a number you can compute from any picture without knowing anything about what is in it. It can tell a human face from a forest, every time, with no training and no idea what either thing is. It cannot tell a painting from a photograph — and that failure turns out to be the interesting part.
The number is the slope of the image’s spatial-frequency power spectrum. In plain terms: take a picture, ask how much of its energy sits at coarse scales versus fine ones, and measure how fast that falls off as you zoom in. Smooth pictures lose energy quickly and have a steep slope. Busy, detailed pictures hold on to it and have a shallow one. That is the whole measurement. It knows nothing about faces, or paint, or art.
The world answers the same thing every time
We measured 101 photographs of natural scenes — forests, coastlines, mountains. They come back at -1.97, with a standard deviation of 0.32. This is not our finding; it is one of the most reproduced results in vision science, and reproducing it is how we knew the pipeline was working before we pointed it at anything interesting.
Before a single image was downloaded, the estimator was run against synthetic fields with a known exponent: white noise, 1/f² and 1/f³. It recovered +0.011, −1.988 and −2.991 — inside 0.012 of the truth in every case. An estimator that cannot recover an answer you already know is not an estimator, and every number below would have been decoration.
A face is not a landscape
Now point the same code at photographs of human faces, cropped so the face fills the frame. They come back at -2.99 — more than a full unit steeper than the world those faces are standing in (t = -10.7).
A face, statistically, is nothing like a landscape. It is far smoother: skin is a broad, low-frequency surface where a forest is a mass of fine structure. So the number separates the two cleanly, and it does so without any notion of what a face is. That is already a slightly uncomfortable result — a two-line calculation is doing categorisation that feels like it should require understanding.
A president nobody ever photographed
So we gave it something it should have no way to handle: a face that was never alive in front of a lens. Gilbert Stuart’s George Washington, painted in 1796, decades before photography existed.
It returns −3.08. The official photographic portrait of Donald Trump returns −2.98. Two presidents, two hundred and thirty years apart, one of them assembled by hand out of pigment by a man looking at another man — and the measurement puts them a tenth of a unit apart.
That is not a coincidence of two pictures. Across the whole corpus, 109 painted faces average -3.13 against 56 photographed ones at -2.99. The difference is -0.14, with a 95% bootstrap confidence interval of [-0.34, +0.06]. It straddles zero.
A null is only a finding if the test had power
“We could not tell them apart” is worthless on its own — you can fail to tell anything apart with a bad enough measurement. So this was written as an equivalence test rather than a difference test: the interval has to straddle zero and the whole interval has to sit inside ±0.35. A gap large enough to matter would have shown up. It didn’t.
The honest reading: painters, working by eye, with no instrument and no knowledge that this statistic exists, reproduced the spectral signature of a real human face closely enough that a measurement built to separate faces from forests cannot separate their work from a photograph.
What we failed to replicate
Two published results did not survive on this corpus, and it would be dishonest to report the ones that did without them.
The first is the well-known claim that paintings in general sit in the natural-scene band — that artists reproduce the statistics of the natural world. Ours do not. Paintings average -2.60 against nature’s -1.97, a gap of 0.62 at t = −14.4, and it is stable when the analysis resolution is halved, so it is not a resampling artefact.
The second is Graham & Field’s specific face result: that painters render faces at natural-scene statistics (~−2) rather than at real-face statistics (~−3), implying they regularise the world rather than copy it. Ours land at −3.13 — on top of real faces, not away from them.
We are not claiming those papers are wrong. Different corpora, different eras of digitisation, different crops. But we measured what we measured.
Why the number cannot rank anything
The temptation here is enormous: a single number, computed from pixels, that tracks something real about pictures. Surely it says something about quality.
It does not, and there is a well-documented case of exactly that mistake. In 1999 a fractal analysis of Jackson Pollock’s drip paintings was published in Nature, and the method was later used to assess the authenticity of disputed works. In 2006, also in Nature, Jones-Smith and Mathur showed the paintings were fractal over too small a range to be meaningful — and that a figure drawn in Photoshop satisfied the same criteria.
The measurement tracks what is depicted, not how well. It separates faces from forests because faces and forests genuinely have different structure. It cannot separate a masterpiece from a doodle because that is not a property it is looking at, and no amount of confidence in the arithmetic changes what the arithmetic is measuring.
What we cannot see, and which way it bends
The face corpora are built with a detector trained on photographs. It finds fewer painted faces than photographed ones, which thins the painted sample toward the most photo-like portraits. That biases against our own finding rather than for it: if anything, the true painted population is further from photographs than we measured, not closer.
The corpus is what the Art Institute of Chicago and Wikimedia Commons have digitised and released, which is not a random sample of all art. Slopes are measured on 512 px square crops through a Hann window — without that window the FFT reads the image border as a step discontinuity and injects its own 1/f², which would have manufactured the headline out of nothing.
And the equivalence is still a null, reported with its interval. It says we could not distinguish two populations with a test that had the power to do so. It does not say they are identical.
One last honest note, because it nearly went the other way. The obvious hero image for this was van Gogh’s Self-Portrait — recognisable in a fraction of a second. It measures −2.23, well off the painted-face median of −3.12. Using it would have illustrated the opposite of what the corpus shows. The pipeline refuses any pinned image sitting more than 0.35 from its category’s centre, which is why you are looking at Stuart’s Washington instead. Fame is not evidence.
Sources: Art Institute of Chicago open access (public domain) · Wikimedia Commons · Graham & Field, Statistical regularities of art images and natural scenes, Spatial Vision 21 (2007) · Redies et al., PLOS ONE (2010) · Jones-Smith & Mathur, Fractal analysis: revisiting Pollock’s drip paintings, Nature 444 (2006). Every figure reproduced from the raw images by data/painting_stats.py.