Workload

long-context-v1

released

Question it answers

How deep is the useful context on this system, not the configurable one?

What is measured

Needle retrieval across a nominal 2k/8k/16k/32k context ladder — 12 items, EN (Dickens) + RU (Dostoevsky); measured prompt tokens run below nominal (EN within ~6%, RU at ~65%) and ship beside the corpus — plus unanswerable controls with a present distractor; format, language, repetition, needle and control gates; quality, TTFT and end-to-end time at each depth.

Claims this workload cannot support

“Handles 128K context” from a memory-allocation fact alone; comprehension claims — this revision varies length at a fixed ~50% needle position, so position sweeps and summarization are future revisions.

Frozen identity

Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.

Revision
2026-08-03
Load model
single_stream
Quality gates
format_gate, language_gate, repetition_gate, needle_gate, control_gate
Metrics contract
quality_at_depth, ttft_ms_at_depth, e2e_ms_at_depth
Corpus hash
2184f558eb79…
← Workload library