Workload

interactive-assistant-v2

released

Question it answers

Is single-user streaming chat responsive on this system for everyday human tasks?

What is measured

Same contract as v1 — client-side TTFT, inter-token latency, end-to-end time; format, language and repetition gates — over a human-task corpus: everyday requests (messages, emails, explanations, plans) instead of self-referential benchmarking questions. 16 items, EN+RU, three length bands.

Claims this workload cannot support

“Supports N employees”, production capacity claims.

Frozen identity

Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.

Revision
2026-08-02
Load model
closed_loop
Quality gates
format_gate, language_gate, repetition_gate
Metrics contract
ttft_ms, inter_token_latency_ms, e2e_ms
Corpus hash
660eaad10d3f…
← Workload library