Column
Essays
Opinion pieces by the lab’s operator: working notes on using AI daily, what breaks, what pays off, and what history says about it. Opinions, not measurements — measured results live in the claim registry.
- We rewrote our benchmark to sound like a human
Our first chat corpus asked the model about build digests and goodput. Nobody's users ask that. The same cell re-measured on sixteen everyday prompts gave a different failure rate on identical hardware, and we kept the old number anyway.
September 2, 2026
- A website where CI fails when a number drifts from its evidence
Type a tokens-per-second figure into this site and CI fails. Every measured value is re-derived from raw run records, and a committed value that drifts from them stops the pipeline. How the gate works, what it enforces, and the class of error it cannot see.
September 2, 2026
- Independent proof or it didn't happen: DeepSeek just made its scoreboard fillable
DeepSeek opened its agent harness and named the mode, effort and sampling behind V4-Pro's agent scores. Better disclosure than the norm, and still not proof: a vendor's number stays vendor-disclosed until a second party runs the same config and publishes the traces.
September 2, 2026
- Agentic coding is a prefill problem, and everyone is shopping for decode
Ask how many tokens per second a local coding agent needs and every answer quotes decode speed. But an agent's turn is a huge prompt and a short reply — the wait lives in prefill, and the hardware conversation is optimizing the wrong half of the request.
September 1, 2026
- Agents-per-box: the capacity metric nobody prints on the box
The agent-swarm conversation runs on single-stream decode numbers, which is the one metric that stops mattering the moment you run a swarm. What we learned sizing shared boxes, and the question buyers should be asking vendors instead.
September 1, 2026
- Everything our harness refuses to measure: a changelog of distrust
A benchmark harness earns trust by what it declines to measure. Every version of ours is a scar with a commit message: the invalid first run we kept, the neighbor check, the fallback detector, the word that matched inside another word.
September 1, 2026
- Every benchmark number is a config in disguise
Two honest people measure the same model on the same hardware and disagree hard. Nobody lied: a published speed is a compressed configuration, and the compression is lossy. What a number must carry to be worth quoting, and how to read one that carries nothing.
August 20, 2026
- The dashboard lied twice in one day, in opposite directions
One box, one afternoon: a model returning empty answers under HTTP 200, and two healthy containers marked unhealthy by a probe aimed at the wrong port. Green-but-broken and red-but-fine are the same disease, and the cure is the same too.
August 20, 2026
- I use AI every day. And I think the Luddites were right
A year of journal entries: every time AI lied to me and every time it genuinely came through. The verification protocol that grew out of it, and why the real Luddite story is about quality and the price of the transition, not fear of machines.
August 14, 2026