Workload

doc-session-v1

released

Question it answers

Does the second question over the same pasted document pay the full prefill again on this system?

What is measured

Client-side time to first token for question pairs sharing a byte-identical document prefix (1 warmup item + 4 documents, 8k/32k tokens, EN+RU), with the server prompt cache off (baseline) and on; format/language/repetition gates on the answers.

Claims this workload cannot support

Cache benefit for non-identical documents; multi-user cache isolation; capacity claims.

Frozen identity

Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.

Revision
2026-08-03
Load model
single_stream
Quality gates
format_gate, language_gate, repetition_gate
Metrics contract
ttft_ms_q1_vs_q2, e2e_ms
Corpus hash
dce620c4d6eb…
← Workload library