Column

We rewrote our benchmark to sound like a human

Our first chat corpus asked the model about build digests and goodput. Nobody's users ask that. The same cell re-measured on sixteen everyday prompts gave a different failure rate on identical hardware, and we kept the old number anyway.

opinion September 2, 2026

This is an opinion column, not a lab result. Numbers here come from the author’s working journal and quoted external sources (marked in the text) — never from the claim registry. Measured, evidence-backed results live at /claims/.

The first version of our chat benchmark was written by an engineer for an engineer, and it reads that way. I know because I wrote it. The workload is interactive-assistant-v1, frozen at the start of August, and its prompts sound like a job interview at this lab. “In one sentence: what does a runtime build digest identify?” “What is the difference between throughput and goodput? Answer in two sentences.” The long-form item wanted an internal note on why a lab should publish failed and invalid runs, with a rebuttal to the objection that it looks unprofessional. We had asked the model to write our own editorials, then timed how long it took to start.

The shape of the corpus was fine. Two languages, three length bands, a word limit per item that the format gate checks, a language gate, a repetition gate, client-side timing from the first token to the last. The problem was the subject. Every prompt tested how comfortable the model was with our jargon, and jargon pulls long, expert answers. Ask a model what a reader still needs before trusting a throughput figure and it writes an essay. Nobody’s users ask a local chatbot what a build digest identifies. They ask it to text the neighbor about the plants.

So the same evening we wrote interactive-assistant-v2. Same shape, same gates, same word limits, and not one prompt about our trade. A two-sentence message asking a neighbor to water the plants this weekend. A polite refusal of a Saturday shift. Why the sky is blue, for a ten-year-old. An email to the landlord about a leaking kitchen tap. In Russian: a parcel for the neighbor to collect from the courier, why winter is cold, a letter to the building management about the stairwell light that has been out for a second week. Nothing a person would be embarrassed to type.

Then we ran the same cell again. Diff the two run manifests and you get the run id, the workload id, the corpus hash and the timestamps. Nothing else moved.

The failure we watch most closely on that cell is the answerless request: the reasoning pass eats the whole completion window, the answer never starts, and the server returns HTTP 200 with nothing in it. On the engineer’s corpus, at the tighter of our two budgets, three requests in four came back empty. On the human corpus, just under six in ten. I had expected the friendlier prompts to make the number worse. It went the other way. The old corpus had been flattering nobody and overstating the failure. In the raw request records the model reasons a little less about household matters than about our trade, and when the window is wide enough to finish, its answers come out at less than half the length. It is the shorter reasoning that gets more answers started before the tight window closes; the shorter answer is a courtesy to the reader, not the reason the request survives. The human-corpus claim pages also fold in later runs on our second, commercially identical unit, which is why they carry a different evidence level from the old ones. On the original box alone the change has the same direction and roughly the same size.

The serving side did not care what the prompt was about. With reasoning off, the first answer token arrived in about a fifth of a second on the engineer’s corpus and on the human one, close enough that I would not bet on the difference. With reasoning on and the roomier budget, the wait was twenty-odd seconds either way, a touch shorter on the everyday prompts than on ours. Hardware and runtime measured the same. The failure rate did not, and the only thing standing between those facts is a text file.

That text file is part of the number’s identity, the same way a firmware version or a parallel-slots default is; an earlier essay already made that argument for every other field. Quote an answerless rate for this model without saying what it was asked and you have quoted half a number.

The obvious shortcut is a public benchmark set. We did not take it. The workload measures timings and gates, not knowledge, and we had just watched the wording of a prompt move how long the model reasons. A prompt the model met in training moves it too, in a direction nobody can see; a set written the evening before the run cannot have been. That is a judgment, not a measurement, and a public set with word limits our gates could check would be worth running beside ours.

We did not fix the old number. Both corpora sit frozen in the catalog with their hashes and in the public harness repository. The two claim families live in the registry side by side, one per workload, on the same system, runtime, model and cells. The catalog validator refuses to let one claim cite runs from two workloads or two revisions, so the v1 and v2 rates can never be averaged into a friendlier figure. The old claim pages stand with their own runs and limitations, and the answerless report cites both, the human number first.

Rewriting a benchmark is a version bump with a diff, not a quiet edit. The temptation is to overwrite the badly written corpus and let the better number bury the worse one in somebody’s old screenshot. Keeping both means anyone can put the two claim pages side by side, check that the runtime fingerprint matches, and decide for themselves what sixteen prompts are worth. That is also the only reason to trust the new number more than the old one: the diff that produced it is public, like the harness’s.

Measured on prompts a person would actually type, the failure is smaller than we had been reporting. At the tighter budget the neighbor gets the text about the plants every time, and the ten-year-old asking why the sky is blue still gets HTTP 200 and an empty message more often than not.

← All columns