Workload
long-context-v1
Question it answers
How deep is the useful context on this system, not the configurable one?
What is measured
Needle retrieval across a nominal 2k/8k/16k/32k context ladder — 12 items, EN (Dickens) + RU (Dostoevsky); measured prompt tokens run below nominal (EN within ~6%, RU at ~65%) and ship beside the corpus — plus unanswerable controls with a present distractor; format, language, repetition, needle and control gates; quality, TTFT and end-to-end time at each depth.
Claims this workload cannot support
“Handles 128K context” from a memory-allocation fact alone; comprehension claims — this revision varies length at a fixed ~50% needle position, so position sweeps and summarization are future revisions.
Frozen identity
Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.
- Revision
- 2026-08-03
- Load model
- single_stream
- Quality gates
- format_gate, language_gate, repetition_gate, needle_gate, control_gate
- Metrics contract
- quality_at_depth, ttft_ms_at_depth, e2e_ms_at_depth
- Corpus hash
- 2184f558eb79…