Column

A website where CI fails when a number drifts from its evidence

Type a tokens-per-second figure into this site and CI fails. Every measured value is re-derived from raw run records, and a committed value that drifts from them stops the pipeline. How the gate works, what it enforces, and the class of error it cannot see.

opinion September 2, 2026

This is an opinion column, not a lab result. Numbers here come from the author’s working journal and quoted external sources (marked in the text) — never from the claim registry. Measured, evidence-backed results live at /claims/.

Type a tokens-per-second figure into a page on this site and CI fails. Put one in the way the rules allow, through a claim id, and get the id wrong, and the static build throws before there is anything to deploy. It was the first design decision of the rebuild.

The rule is that a measured value reaches a page only through a component that takes a claim id; the scanner enforces the throughput-shaped part of that rule, and the rest of it is enforced by me. An unknown id throws, with a message telling you to add a claim backed by valid runs first. If the claim exists but has been superseded or retracted, it throws too, because a replaced number has no business circulating in a sentence. What renders is the value, its unit, and a link to a permanent page carrying the claim’s scope, limitations and runs.

The registry is generated, never edited. Upstream of it a claim is a JSON file I do write by hand: a statement in words, the run ids it rests on, a path to a SQL file, its scope, limitations, evidence level and status. The value is the one field that is not there. A script runs each claim’s query with DuckDB over the per-request records of exactly the runs the claim names, looks up which machine and model artifact those runs used, and writes the result into one generated TypeScript file whose header says do not edit. CI runs the same script in check mode and fails when the committed file differs from what it just derived. Edit a value by hand, or touch a run without re-deriving, and the pipeline stops. The registry also ships as JSON and as a markdown twin per claim.

I did not build it this way for elegance. The rebuild had one job: a number must not be able to outlive its evidence. The registry was born empty and stayed empty for days after the site went live; the pages meant to hold numbers showed empty states, and the comment at the top of the registry file still says legacy numbers must not be re-added. Today the registry holds a few dozen claims. I wrote every statement, and I have not typed one value.

What CI enforces

The denominator lives in the SQL, not in the methodology page. The query behind the answerless-response claim carries a comment: the denominator is every request issued in the referenced runs, failures are never dropped. The SELECT under the comment does exactly that: it divides by a bare count over every record in the runs the claim names.

The validator has to prove it can say no, for the same reason the harness keeps its invalid first run. Next to the catalog there is a folder of mini-catalogs built to be rejected: a claim active on a run still marked planned, a claim spanning two workloads, a claim aggregating across systems without a comparison class, a model with no hash, a run whose references point at nothing. The catalog gate fails if it accepts any of them. A validator that has never been shown rejecting something is a formatting step with a confident name.

Historical quotes need a marker on the line. The wording scanner refuses a hand-typed throughput figure in page copy unless the line carries a small HTML comment meaning “reviewed quote from a named source”. Numbers from other people’s publications may appear; they may not look like ours.

For the first few days the drift check lived only in my local pre-push script, which means it ran when I remembered. Now it runs in CI, on a machine that is not mine.

What green CI does not mean

It does not mean the sentences are true. Not one entry in our errata log was caught by the gate.

The clearest case is the Gemma control-success claim. Its statement read “admitted the answer was absent instead of fabricating one”. The archived per-request records show the failed controls were empty answers: the reasoning pass ate the token budget before any text came out. No run fabricated a code. The value never moved. The statement around it was wrong, and wrong against the model, and the pipeline could not tell, because it checks arithmetic and the error was in English. The fix was an errata entry and a rewording, linked both ways. As I write this, the comment at the top of the SQL file that computes that value still says “instead of fabricating one”. CI does not read comments either.

Then the long-context ladder, whose nominal token counts came from a characters-per-token estimate. Measured against the runtime’s own accounting, English rungs landed near nominal and Russian ones far below it. The corpus bytes were untouched and still matched their frozen hash. The error was in describing the evidence, not in the evidence, and a hash cannot see that. A re-read caught it; measured per-item counts now ship beside the corpus.

Then a word: site copy said “preregistered” before any preregistration artifact had been frozen. Reworded after a review.

And one more. The morning the drift gate went remote, the errata page was updated to say no corrections had been needed yet. The sentence survived part of one morning. The disclosures page kept saying no corrections had been recorded for days after the log had several. The page about corrections needed correcting.

The gate checks the arithmetic between a number and the runs it names, nothing else, and I would build it again anyway. Arithmetic is the part that gets edited quietly at midnight, when a result looks better with one run dropped and nobody would notice. Sentences get edited in daylight, in front of everyone, and still come out wrong. The numbers are guarded by a script that cannot be talked into anything. The sentences are guarded by me, and I can be.

← All columns