Workload
long-context-v1
Draft
Question it answers
How deep is the useful context on this system — not the configurable one?
What is measured
Depth ladder 8K/32K/64K/128K; RULER-derived tasks plus LongBench-style subset with answerable/unanswerable controls; distinguishes configured / loadable / completed / quality-preserving / recommended.
Claims this workload cannot support
“Handles 128K context” from a memory-allocation fact alone.
Identity discipline
Every released workload revision freezes its item hashes, tokenizer revision, load model, cache policy and quality gates. Results reference the exact revision; changing any identity field creates a new revision.