Program

Reproduce a claim

The highest rung of this lab's evidence ladder — external_reproduced — is empty. That is deliberate: it can only be filled by someone who is not us. If you own comparable hardware, this page is the standing contract.

The contract

  1. You run the reproduction kit for a listed claim on your hardware: pinned runtime image digest, hash-verified model artifact, frozen corpus, one command. Kit v1 covers gate-based claims — needle retrieval and unanswerable-control honesty — where the tolerance is binary: the same gate outcomes on your unit, or not.
  2. You submit the sealed evidence bundle the harness produces (a GitHub issue in agmind-bench with the bundle attached or linked).
  3. We verify the bundle against the tolerance file. Match → the claim page links your bundle, the claim's evidence level rises to external_reproduced, and you are listed here with a permanent link to whatever you choose: your repo, your write-up, your handle.
  4. No match → that is a finding, not a rejection. We publish the divergence and investigate the configuration delta in the open, the way our errata page already works. Failed reproductions are first-class results.

Why bother

You get a permanent, linkable artifact proving your hardware's behavior under a frozen, documented workload — measured by an instrument with public unit tests, not a vibes benchmark. We get the only thing we cannot produce ourselves: independence.

The kit

Everything is in the public harness repository: the runner, the corpora frozen by hash, the artifact fetcher with published checksums, and the tolerance file naming the claims open for reproduction.

agmind-bench → REPRODUCE.md → Start here: /claims/

Roster

No external reproductions yet. This section stays honest: it fills when the first bundle passes verification — or documents the first divergence.