AGmind Request a scope

Workload

structured-agent-v1

Draft

Question it answers

Does JSON/function calling stay reliable under realistic agent traffic?

What is measured

BFCL-style calls including relevance rejection, parallel and executable calls, loop detection; deterministic sandbox; parsing success alone is not task success.

Claims this workload cannot support

Agent-readiness claims from parse rate alone.

Identity discipline

Every released workload revision freezes its item hashes, tokenizer revision, load model, cache policy and quality gates. Results reference the exact revision; changing any identity field creates a new revision.

← Workload library