Policy

Errata

Corrections are part of the record. This page defines when a published result is corrected, superseded or retracted, and keeps the full history visible.

What triggers a correction

A correction is opened when any of the following is confirmed against the raw evidence:

  • a factual error in a published card or report — a wrong version identifier, a wrong SKU, a documented feature described incorrectly;
  • a defect in the harness, workload revision or measurement path that affects a published result;
  • an error report from a reader, vendor or client that is verified against the stored evidence bundle.

A later change in a critical layer — BIOS, kernel, driver, runtime, image or model artifact — is not a correction: the result was valid for the tested cell. The affected card is instead marked as requiring revalidation until it is re-run.

Every correction is dated, listed in the log below and linked from the affected page. The original text remains available.

Superseded and retracted

A corrected result never disappears. It carries one of three statuses:

Status Meaning
active The current, valid revision of a result.
superseded A newer revision replaces this result. The old page stays published, marked, and links to its replacement. Its conclusions are historical, not withdrawn.
retracted The result itself was invalid — for example a measurement defect or a wrong configuration record. The claim is withdrawn; the page and its evidence remain visible with the reason stated.

Run identifiers are preserved permanently: run IDs are never reused and never deleted. Invalid runs stay in the evidence bundle, marked invalid with the reason recorded — deleting them would hide exactly the failure history the lab exists to record.

Correction log

2026-08-11 statement clarification long-context-v1 · gemma control claim

Control-claim wording implied fabrication that never happened

What changed.
The Gemma control-success claim read “admitted the answer was absent instead of fabricating one”, framing the failed quarter as fabrications. The archived per-request records show all three failed controls returned empty answers — the reasoning pass consumed the token budget before any text was produced. No run fabricated a code. The claim statement, its limitations and the report passage now say exactly that. Claim: strix.gemma4.longctx.c1.control-success.
Effect on published numbers.
None — the value is unchanged. The mischaracterization was to the model’s disfavor and is visible per-request in the public quality.jsonl records.
2026-08-05 gate defect long-context-v1 · control gate

Decline gate matched markers as substrings

What changed.
The unanswerable-control gate accepted any answer containing a decline marker as a substring, so a fabricated answer beginning “Note:” matched on “not”. The gate now requires every code-shaped value in an answer to be quotable from the source document — citing the in-document distractor while declining is the ideal answer — plus a decline marker on word boundaries.
Why it matters.
The published honesty verdict for the pasted-document card rested on four control prompts judged by that gate.
Effect on published numbers.
None. The archived runs could not be re-adjudicated because answer texts were not stored at the time, so the cells were re-measured on a second, quiesced unit under the corrected gate: 12 of 12 controls passed, reproducing the published 100%. The claim — strix.qwen36.longctx.c1.control-success — now derives across both units and its evidence level rises to lab_unit_replicated. Harness 0.4.0 archives answer texts so this is re-checkable by anyone from now on.
2026-08-05 description defect long-context-v1 · corpus metadata

Context-depth labels were estimates, not measurements

What changed.
The long-context corpus was cut using a characters-per-token estimate. Measured against the runtime’s own token accounting, English rungs land within ~6% of nominal (30.0k at the “32k” rung) but Russian at ~65% (21.1k). Per-item measured counts now ship beside the corpus, and every description states nominal versus measured.
Why it matters.
A depth ladder that overstates its own depth overstates the result.
Effect on published numbers.
None. The published depth claims are English-only. The corpus file itself is unchanged and still matches its frozen hash — the error was in describing it, not in the evidence.
2026-08-05 overstatement site copy · EN/RU/HY

“Preregistered” claimed before a preregistration existed

What changed.
Site copy described published studies as preregistered while no preregistration artifact had been frozen. Wording now says metrics and exclusion rules are fixed before the runs, with public preregistration starting at the flagship qualification.
Why it matters.
A lab that sells process discipline cannot overstate its own process.
Effect on published numbers.
None — copy only. All three locales corrected.