Column

Independent proof or it didn't happen: DeepSeek just made its scoreboard fillable

DeepSeek opened its agent harness and named the mode, effort and sampling behind V4-Pro's agent scores. Better disclosure than the norm, and still not proof: a vendor's number stays vendor-disclosed until a second party runs the same config and publishes the traces.

opinion September 2, 2026

This is an opinion column, not a lab result. Numbers here come from the author’s working journal and quoted external sources (marked in the text) — never from the claim registry. Measured, evidence-backed results live at /claims/.

For months you could argue with a frontier agent score but you could not run it. The model was open, sometimes; the benchmark was public, usually; the scaffold that turned the model into an agent was the vendor’s private tool, always. In August that excuse ended for one lab. DeepSeek shipped the release build of V4-Pro under an MIT license, opened the agent harness it scores with, also MIT, and wrote on the model card how the code-agent rows were produced: “For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.”

That sentence says more than most cards do, and I want that on the record before the rest. Our own planning note filed this essay under “V4-Pro’s empty scoreboard — vendor benchmarks measured in the vendor’s own unreleased harness.” That line did not survive August, and I am glad to retire it. Framework named, mode named, the highest effort level named, sampling named. The usual scoreboard gives you a bold number and a footnote that says “internal”.

It is still not proof, and the reason has nothing to do with honesty. A scoreboard row is a configuration, and a number is a config in disguise. Write down what one agent-benchmark row depends on. The harness commit: the repo is a developer preview that promises breaking changes, so “DeepSeek Harness” without a hash names a moving target. The plugin set, since the architecture is everything-is-a-plugin and what “minimal mode” includes is a definition the harness owns, not the card. The tool sandbox the agent ran inside. The effort level and the sampling, which the card does give. The serving stack under all of it: runtime build, quantization, and the speculative-decoding module they call DSpark, which has a version too. The benchmark’s own version. And the raw traces, request by request, that every one of those choices left behind. The card publishes the top of that list. Nobody can re-derive a number from a list. We looked for published traces behind the model-card table and could not find any; the repo carries a BENCHMARK.md I have not yet worked through, so read that as a question, not a verdict.

I am not being theoretical about the serving stack. We serve the smaller sibling, V4-Flash, across a pair of DGX Sparks, and “the model” changed under us without a weight changing. I have told that before: a kernel flag the automatic mode skips, where the only way to know which tier you were in was a line in the server log, and a set of published speeds for that model on those two boxes, each of which could be right once pinned to its own settings. The incident I have not told is from one of those boxes running a different MoE model: a quantization path was broken in the mainline build we tested, and the checkpoint with half the bytes decoded slower than its FP8 sibling. A scoreboard row inherits all of this. Quality rows inherit it more quietly, because a wrong kernel announces itself in throughput and a wrong sampling default announces itself in nothing.

Every card and report we publish carries an evidence level, a ladder for how much measurement stands behind a number: one valid run, repeats on the same unit, a second identical unit, reproduced outside the lab. Off to the side, not on the ladder, sits vendor_disclosed. A vendor’s number stays there forever. The vendor can repeat the run as many times as it likes and it stays there, because repetition by the same party with the same incentive is one data point in a larger font. It climbs when someone else, with a different reason to care, runs the same configuration and publishes the traces. Until this August that was structurally impossible for this model family. Now it is merely unpaid work. Independent evaluators exist: Artificial Analysis lists the model at max effort on its Intelligence Index, a versioned composite of several evaluations. Good. A composite index is a configuration too, chosen by a third party with its own scaffold, so it is a different row and not a reproduction of the vendor’s. The vendor’s row, in the vendor’s harness at a named commit with traces attached, is still empty. Whoever fills it first owns that number.

Now the standard, applied to us. We cannot run V4-Pro. Our pair holds the Flash sibling with tensor parallelism across both nodes, and the Pro would need many times that memory; the honest version of “we evaluated it” for this lab is “we did not”. The DGX lane is also bring-up work: what those pages ship is a public run repository with a manifest and raw outputs, plus a note on each that none of it has reached the claim registry. The row we can fill in full is the registry lane, with the bundle anyone needs to re-derive a number. Every run there leaves a manifest, the per-request records with the failures still in them, the server log, and checksums over the lot. The image digest, the model revision and the corpus hash sit in the manifest next to the settings, and the host inventory is taken at start and again at end, so quiescence is a check that can fail, not a sentence in the notes. The published numbers are re-derived from those records in CI, and the standing offer is that anyone who runs the kit for a listed claim and sends a bundle that matches lifts it to external_reproduced, while a bundle that does not match gets published as a divergence and investigated in the open.

An open harness turns “trust us” into “run it”. Somebody with different incentives and enough memory still has to type the command, against a commit the card does not name. Until they do, the scoreboard is a promise with excellent documentation.

← All columns