Workload
structured-agent-v1
Question it answers
Does strict JSON output stay reliable for everyday automation tasks on this system?
What is measured
16 deterministic items, EN+RU, three task types — extraction with exact ground truth, classification with closed label sets, generation under a structural contract. Strict JSON parse plus per-key match separate parse success from task success; the repetition gate approximates loop detection. No tool-execution sandbox in this revision.
Claims this workload cannot support
Agent-readiness claims from parse rate alone; tool-calling claims — this revision executes no tools.
Frozen identity
Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.
- Revision
- 2026-08-03
- Load model
- closed_loop
- Quality gates
- format_gate, repetition_gate, json_gate
- Metrics contract
- task_success_rate, parse_success_rate, loop_detection_rate
- Corpus hash
- 9687fb0512b4…