Column

Everything our harness refuses to measure: a changelog of distrust

A benchmark harness earns trust by what it declines to measure. Every version of ours is a scar with a commit message: the invalid first run we kept, the neighbor check, the fallback detector, the word that matched inside another word.

opinion September 1, 2026

This is an opinion column, not a lab result. Numbers here come from the author’s working journal and quoted external sources (marked in the text) — never from the claim registry. Measured, evidence-backed results live at /claims/.

The first run in our public archive is invalid. Not deleted, not rerun and quietly replaced — invalid, marked as such, kept on purpose. The harness was version 0.1.0, the go/no-go pass produced no output worth keeping, and the record of that failure is the first brick of everything we published since. I have come to think of the harness’s whole changelog that way: not a list of features, but a list of things we stopped trusting, one incident at a time.

We stopped trusting the host. Early on, a measurement ran while something else woke up on the box, and the numbers wobbled for reasons that had nothing to do with the model. Now every run records a host inventory twice — once at start, once at end — and a GPU-capable neighbor appearing in either snapshot invalidates the cell. Not flags it. Invalidates it. The two-node fleet exists so one box can stay empty enough for that check to pass: the dirty box runs everything else precisely so the clean box can run nothing.

We stopped trusting the runtime to fail loudly. A model can quietly leave the GPU and keep answering — slower, correct, and invisible unless you go looking. The harness now reads the server log for fallback markers, because a silent CPU fallback produces beautifully reproducible numbers about the wrong device. This is the measurement cousin of the green-but-broken dashboard: the worst failures are the ones that keep producing output.

We stopped trusting our own plumbing. One endurance run died not from the model but from the harness itself: a thread-pool convenience method eagerly queued the entire workload’s futures, the queue ate the machine’s memory, and the out-of-memory killer ended a three-hour measurement in hour two. The fix is a comment in the source now, in capital letters, so the next person — usually me — does not rediscover the convenience.

We stopped trusting substrings. The gate that checks whether a model honestly says “that’s not in the document” once matched the marker not inside the word Note, and a fabricated answer sailed through as an honest refusal. The fix was word boundaries, then a harder rule behind it: any code-shaped value in an answer must be literally quotable from the source document, or the answer fails. The decline-gate report stands on that fix; without it, the gate was measuring our regex, not the model.

And underneath all of it, the rule that predates the harness: failures stay in the denominator. Every dropped request, every empty answer, every invalid run is part of the result, because the alternative — averaging only the requests that behaved — turns a reliability problem into a speed bonus and rewards the server that fails fastest.

None of these checks made our numbers bigger. Several made them smaller, and one deleted a result I liked. That is roughly the point. A harness that only ever confirms is not measuring; the version history of your distrust is the most honest documentation your numbers have. Ours is public, scars and capital-letter comments included — partly so you can run it, mostly so you can check what it refuses to do.

← All columns