Workload

structured-agent-v1

released

Question it answers

Does strict JSON output stay reliable for everyday automation tasks on this system?

What is measured

16 deterministic items, EN+RU, three task types — extraction with exact ground truth, classification with closed label sets, generation under a structural contract. Strict JSON parse plus per-key match separate parse success from task success; the repetition gate approximates loop detection. No tool-execution sandbox in this revision.

Claims this workload cannot support

Agent-readiness claims from parse rate alone; tool-calling claims — this revision executes no tools.

Frozen identity

Every released workload revision freezes its item hashes, load model and quality gates. Results reference the exact revision; changing any identity field creates a new revision. The facts below are read from the catalog file, not typed into this page.

Revision
2026-08-03
Load model
closed_loop
Quality gates
format_gate, repetition_gate, json_gate
Metrics contract
task_success_rate, parse_success_rate, loop_detection_rate
Corpus hash
9687fb0512b4…
← Workload library