reelgorithm$39.99

Chapter 7 of 123,221 wordsFree, and no email

Things that cost a day

This is one complete chapter from Anatomy of a Data Reel, published in full. Every failure below actually happened, cost a day or more, and passed every check that was in place at the time — which is the only reason any of them are worth writing down.

This is the free chapter. It's the one I'd read first if I were deciding whether to buy the rest.

Every failure in this chapter has one property in common, and it's the property that makes them expensive: they exit zero. Nothing crashes. No error is printed. The file gets written, the render completes, the gate goes green, and the thing is wrong.

That's not a coincidence, it's the whole category. A crash costs you twenty minutes. A silent wrong answer costs you a day, or it ships. I've lost more time to tools that returned a plausible number without doing the work than to every tool that ever threw an exception.

Here are the ones that got me.


1 · Four ways a GPU scene renders pure black

A single-frame renderer is not a browser. Anything in a WebGL stack that assumes a continuous animation loop will produce a black rectangle and no error, because from the renderer's point of view a black rectangle is a perfectly valid frame.

The backend. You have to force ANGLE. The default backend renders GL panels black without erroring; EGL and SwiftShader crash outright, which is at least honest. Pass --gl=angle (or swangle) on every render that touches WebGL, and put it in the render script rather than in your memory, because you will forget it once and lose an afternoon deciding your shader is broken.

Instanced meshes from the convenience layer. The helper component for instancing syncs its instance matrix inside the per-frame animation hook. A one-shot renderer never ticks that hook, so the matrix stays at identity and nothing draws. Build the instanced mesh imperatively and write the matrices in a layout effect instead. Same output, and it survives a single-frame capture.

Map layers that render asynchronously. A GPU map library that draws on its own schedule needs an explicit handshake: turn off its internal animation loop and open a render-delay handle per frame, released by the library's own after-render callback. Skip it and stills still pass, which is the cruel part. The full render tears instead, because a still gives the library enough wall-clock slack to finish and a 3,000-frame render doesn't.

Colour strings. Palette helpers that blend or ramp colours usually return rgb(...) strings, not hex. A GPU layer that parses hex only gives you NaN, and then uploads the NaN to the GPU without complaining. Black.

Four traps, four days. And a fifth I'll throw in free: WebGL ignores line width on the basic line material. Raw line segments are always a one-pixel hairline no matter what you set. Use the camera-facing line implementation instead, which honours the width, and which, unlike the instancing helper, sets its geometry on mount rather than in the animation loop, so it survives the one-shot render.

Two more, if you're loading real 3D models. An async model loader will miss a single-frame capture entirely and hand you a black frame; bake whatever geometry you need to a static file offline and import it, which also makes the scene a pure function of the frame number. And a model file whose texture images are missing will abort the whole parse, not just the textures — if you only need the geometry, strip the image, texture and sampler blocks out of the model header before loading it.

How to survive this class permanently: keep two smoke-test scenes, one per GPU library, and re-render both after every dependency bump. Two stills is a cheap standing cost against four days you've already paid once.


2 · The 44 milliseconds that make every sound effect land late

Every sound effect in a video was landing 45–48 milliseconds after its intended frame. Uniformly. Not drifting, not per-cue — every single one, the same amount late.

It reads as "the sound design feels a bit off" and it's almost impossible to find by listening, because 45ms is right at the threshold where audio-late becomes perceptible, and it's worst exactly where it hurts most: sharp transient cues keyed to a visual cut.

It isn't a composition bug. Rendering the identical composition to ProRes with 16-bit PCM audio put the first transient at 5.4ms — exactly the sample's own peak offset, which is to say perfectly in sync.

The lag is AAC encoder priming. The encoder needs 2,112 samples of runway, which at 48kHz is 44 milliseconds, and the muxer writes them into the file without an edit list, so nothing downstream knows to compensate. Probe the file and both streams report start_time=0.000000, which is how you confirm it: the container is asserting they're aligned and they aren't.

The fix is two commands instead of one. Render a PCM master, then encode delivery separately:

# master: ProRes video, PCM audio, no AAC anywhere near it
<renderer> render <id> out/<id>_pcm.mov --codec=prores --audio-codec=pcm-16

# delivery: ffmpeg writes the edit list, so the offset is compensated
ffmpeg -y -i out/<id>_pcm.mov -c:v libx264 -crf 16 -pix_fmt yuv420p \
  -c:a aac -b:a 192k -movflags +faststart out/<id>.mp4

That measures 5.4ms. In sync.

How to verify any file, rather than trusting me: decode the audio to raw mono PCM at a low sample rate, walk the samples until one exceeds a noise floor, and divide by the sample rate to get the time of the first energy. Compare it to the cue time you intended. A uniform offset across every cue means priming. A growing offset means drift, which is a different bug.

What not to do: don't fix this by moving your cues earlier in the timeline. It breaks the editor preview, and it rots silently the day someone changes the encoder.


3 · Two fixes that were each correct, and collided

This one is my favourite, because both changes were right, the bug only exists in the file that gets uploaded, and it is invisible in the editor, invisible in the master, and invisible in every still.

Fix one. The renderer defaults to a BT.601 colour space, so every master was being tagged as standard-definition PAL while the content was sRGB. Decoded as-tagged versus as-BT.709, the brand pink moved about 30 levels out of 255 in two channels. The fix is --color-space=bt709. Correct.

Fix two. Masters were being rendered at twice the platform's ceiling on each axis (four times the pixels), so the platform's fast server-side scaler ran on every frame, and a cheap downscaler destroys thin rules and small type first. The fix is to downscale locally with a good filter to exactly the target resolution so their scaler never runs. Also correct.

The collision. The delivery step had in_range=full hard-coded, and that was right for every master made before fix one existed, because the renderer's default output is full-range. But --color-space=bt709 makes it emit limited-range instead. So fix one turned fix two into a bug: the scaler squeezed an already-limited master a second time.

luma 8  →  16 + 8 × (219/255)  ≈  23

Measured on one episode: the master's background sat at luma 8. The delivered file's background sat at 23. A uniform, washed-out grey where the design says near-black. The report I got was simply "the background isn't the right black" — from someone watching the uploaded version on a phone.

Why this look is maximally exposed to it: about 42% of a frame sits below luma 16 and about 12% above 235. Fifty-five percent of the picture lives at the extremes, because it's white type on near-black. Any stage that disagrees about range mangles over half the frame.

The fix is to detect rather than assume. The delivery step now probes the source's pixel format and tags, decides full or limited from that, and — this is the part worth stealing — deletes its own output file rather than ship one whose tags came out wrong. A tool that refuses to leave a bad artifact on disk is worth more than a tool that warns.

And one specific ffmpeg landmine inside the same fix: passing -x264-params at all overrides ffmpeg's own -color_primaries and -color_trc, and they come out unknown. The command exits zero. Set them inside the params string instead — colorprim=bt709:transfer=bt709:colormatrix=bt709.

How to check it without eyeballing: measure the always-empty margins of the frame — the bands the platform's UI covers, which are guaranteed to contain no content, and compare their luma between master and delivery. Two cleverer metrics were tried first and both returned confidently wrong answers: a "luminance haze band" scores antialiased type, and "smooth dark pixels" scores flat chart fills.


4 · A hockey trophy on a bank

Matching a named entity to an image from a catalogue by fuzzy search is a problem that looks solved and isn't. The rule in place was "the candidate shares at least one distinctive token with the subject". Every one of these matched confidently, downloaded cleanly, and rendered:

wantedgot
Citizens BankCitizens First Bank — a different bank
The Huntington National BankHuntington Heroes — its charity programme
Capital OneCapital One Hall — a concert hall
Glacier BankGlacier County, Montana
Centennial BankCentennial Cup — a junior hockey trophy
First National Bank of PennsylvaniaFirst Bank of the United States — closed 1811
Santander BankSantander Consumer Bank
First Horizon BankFirst Tennessee Bank — its own former name

Look at what they have in common. Every one shares a word and adds another — HALL, HEROES, CUP, COUNTY, CONSUMER, and the added word is precisely what makes it a different thing.

The rule that works: require the candidate's identifying words to be a subset of the subject's, after stripping generic vocabulary and bare numbers. Accept a strong substring match separately. Be deliberately stricter than you need to be, and recover the good ones through a hand-checked pin table.

A subject with no image can fall back to a name tile, which is honest and looks deliberate. A subject wearing another company's mark is a factual error on screen that no downstream check can catch. Never loosen the rule to rescue one asset.

The mirror failure, from the same build: over-aggressive generic-word stripping emptied "Bank of America" to a token-less string, so it matched nothing. And an empty normalised key isn't a failed match — it's a key that matches every other name that also emptied out.

The same class, in brand names. Sourcing airline liveries, a bare /united/i matched China United Airlines — a red tail that shipped into a render as United, and later a British operator, where the match was on the word Kingdom. /american/i matched American Trans Air, a carrier that no longer exists. Match the full name as a phrase, carry an explicit exclusion pattern per brand, and reject any title that names two of your brands at once. That last guard alone catches most of it.

And then look at the images. Every failure above is obvious in one glance and invisible to every automated check — the file downloads, the licence is clean, the aspect ratio is fine, and the plane is the wrong colour. Contact-sheet the whole asset folder before wiring any of it into a scene.


5 · The axis helper that rounds your range and lies about it

A charting library's linear scale usually "nices" the domain by default — rounding the ends outward to friendly numbers. That's a sensible default for a computed domain, and it is a catastrophe for an explicit one.

An explicit domain of [1, 9], for a nine-item rank chart, rendered an axis running 0 → 10. Ticks for rank 0 and rank 10, neither of which exists. No error. The chart looks plausible and the numbers are wrong. The same bug was sitting in an already-published episode.

Two halves are needed to fix it, and doing only one leaves it broken:

  • pass the flag that disables nicing when the caller supplied a domain, so the domain is honoured;
  • filter the generated ticks to those inside the domain, with a small float tolerance, because the tick generator will still emit ones outside it.

Why it matters more than it looks like it does: for an ordinal quantity — ranks, counts, positions — a niced domain isn't a cosmetic difference. It invents categories that don't exist. That's a lie-factor problem, not a layout one.

Fix it in the shared chart code, not in the scene. A scene-level workaround leaves the next video to rediscover it.

So the check is mechanical now. Hand a library an explicit domain, then pull a still and read the axis end labels back against what you asked for. An explicit domain is a request rather than a guarantee, and nothing tells you when the library declined it.


6 · Ranking photos by shape gets you paintings

Sourcing 100 US city photographs from a free image repository, the ranking function preferred portrait-shaped images, because the video is portrait. Three distinct failures, each of which shipped into a render before being caught:

  1. Preferring portrait → paintings and archives. Portrait-shaped "city" content in a free repository is art and documents. Columbus came back as a Childe Hassam oil painting. Austin came back as the cover of a 1976 Armed Forces Week magazine.
  2. Preferring portrait photographs → single buildings. Kansas City became a lobby entrance. Pittsburgh became one tower.
  3. Not filtering orbit → satellite imagery. Houston came back as a photo taken from the International Space Station.

Never rank candidates by aspect ratio. Hard-filter to modern photographs, score for whole-subject views, and use aspect only as a tie-breaker. Reject on title and category words for artwork, print media, maps, monuments and diagrams; reject orbit words; reject single-structure words unless the title also matches a whole-city word; reward skyline, cityscape, panorama, aerial, downtown; require the subject's own name in the title, because search returns neighbours.

Then the framing problem, which is separate and worse. object-fit: cover destroys a landscape photograph in a portrait frame. A 1.5-aspect skyline in a 1080×1920 frame scales to 2880×1920 — you see the middle 37% of the width. No amount of re-picking photos fixes it, because the crop is the problem.

Size the image off width — around 1.4 to 1.55 times the frame width — let the height fall out of the aspect ratio, and anchor it high so the subject sits in the top half. Feather both ends with a gradient mask, or a hard-edged band reads as a photo pasted onto black. Give type its own contrast with a shadow rather than turning the scrim up until the photo dies.


7 · The probe that can't detect the thing it's measuring

Four measurements in one session were wrong in a way that looked like a result. This is the most dangerous item in the chapter and it's the reason the rest of it exists.

  • A simulation of highlight/shadow crush returned 0.08%. The scaler used the frame's own metadata and silently ignored the range override that was the entire condition being tested, so the filter never applied it. Recomputed as plain arithmetic on the luma histogram: 42.4% at or below 16, 12.5% at or above 235. Off by three orders of magnitude, and 0.08% would have ended the investigation at the wrong conclusion.
  • A PSNR comparison emitted nothing. Twice. No error, no output.
  • "Distinct dark tones — master 1, delivery 1." The sampled patch is flat in the master too, so the probe couldn't have returned anything else. It demonstrated nothing while reading as a finding.
  • An encode that "succeeded" shipped untagged — the -x264-params override from §3, exit code zero.

The rule, and it's the one worth the price of the chapter: positive-control your instrument before you believe a null. Feed the probe an input where the effect is known to be enormous. If it doesn't light up, the probe is broken, not the hypothesis.

Three corollaries:

  • Prefer arithmetic on raw values over a filter chain that might reinterpret them. Histograms don't have opinions about colour metadata.
  • Assume your tooling may silently override a flag you passed. Probe the output and assert the properties you asked for. Deleting your own bad output is the pattern to copy.
  • When a measurement is unreliable, say "inconclusive, here's why" and move on. Don't keep chasing it, and never soften it into a weak finding.

8 · The render you're looking at is the old one

A short list, because each of these individually cost an hour and collectively cost more than any single bug above.

Check the output file's modification time after every render. Not the exit code — the mtime, and the duration. A stale mtime means the render failed silently and you are about to QA yesterday's file. I once spent many turns extracting frames from a stale video while two separate failures were being swallowed by a pipe into grep.

Don't pipe build output through a filter that hides errors. If you must filter, also grep for ERROR, Traceback and command not found.

A glob that matches nothing can kill an entire command chain. In zsh, rm -f out/Thing*.mkv && <render> with no matching file aborts at the glob, so the render never runs and the failure looks like an ordinary empty result. Use a literal path, or enable null-glob, and never let a possibly-empty glob gate a long command.

Caches keyed by scene name collide across projects. A text-to-speech cache keyed on scene name meant one video silently played another video's narration in two scenes, because both projects had a scene called ColdOpen. Diagnose by comparing the cached file's timestamp against the script's — stale audio predates the script it's supposed to be reading.

Some verbosity flags suppress the output you're asking for. Running a volume measurement at error-level verbosity hides its report entirely and looks exactly like silence, which is how a perfectly good mix gets diagnosed as a missing audio file.

Check every beat, not just the ones you changed. Two collisions shipped in an already-"verified" render — an annotation drawn over seven rows of a bar chart, and a line of text overprinting an axis, because the earlier pass sampled the beats that had been edited. A full render is minutes. A still is seconds. Every defect found by rendering is a defect a still would have found for about 3% of the cost, and a defect not found ships.


The pattern

Read back through the eight. Not one of them announced itself.

The GPU renders black and reports success. The audio is 44ms late and the container claims both streams start at zero. The delivered file is grey and the master is perfect. The logo is the wrong company and the licence check passes. The axis is wrong and the chart is beautiful. The photo is a painting and the download is clean. The probe returns 0.08% and exits zero. The render is stale and the exit code is fine.

The danger is the absence of an error, not the presence of one. Every gate you build should be designed to fail loudly on the things that otherwise fail quietly, and the single most valuable one you can write is the tool that deletes its own output rather than hand you something wrong.

There's a checklist version of this chapter in the artifact pack: the pre-ship trap list.

That was one of 12.

The rest covers where the data comes from when nobody has published it, how to know your numbers aren’t quietly wrong, and why a chart can be accurate and still lose the argument. 51,134 words, 7 format blueprints and 10 working sheets.

One payment · lifetime updates · 30 days to change your mind