Watch the local-agent conversation for a week and a pattern emerges. Someone asks what hardware they need to run a team of coding agents. Someone answers with a tokens-per-second figure — single stream, one chat, decode only. The question was about a swarm; the answer was about a soloist.
I understand why. Single-stream decode is the number every review prints, every spec-sheet comparison ranks by, every forum thread argues about. It is also close to useless for the question being asked, because an agent swarm is not one fast conversation. It is many slow ones happening at once, each mostly waiting on tools, each firing bursts of work at unpredictable moments. The metric that decides whether your box survives that is not how fast one stream decodes. It is how the wait grows as streams stack up, and where it stops being tolerable.
We measured that growth on our own hardware — one, four, eight simultaneous streams on a 128 GB box — and the shape surprised me more than the values. The wait does not grow linearly. It stays flat long enough to lull you, then bends. Four concurrent chats felt indistinguishable from one; eight was a different machine. The numbers and their limits are in the concurrency report, and the same story at higher stakes — a 284B model aggregating a dozen streams across two DGX Sparks — is in the buyer’s answer. What I want to argue here is not the values. It is the metric.
There is a physical reason agents-per-box is the right unit for this hardware generation. Mixture-of-experts models leave memory bandwidth on the table at a single stream — the bus that decode starves on is not saturated by one conversation. Concurrency is how you buy that headroom back. A box that looks mediocre in a single-stream review can be quietly excellent at four streams, and two boxes with identical solo numbers can diverge completely under load. The soloist benchmark cannot see any of this. It is not wrong; it is answering a different question.
So here is the question I think buyers should ask instead, phrased so a measurement can answer it: how many concurrent streams does this box hold before the first visible word takes longer than my users tolerate? Pick your own threshold — half a second, two seconds, whatever your agents’ feedback loop survives. Then demand the ladder, not the point: the wait at one, at four, at eight, with completion rates attached, because a box that “handles” eight streams by silently dropping every fifth request is not handling anything.
Nobody prints this on the box. Vendors print peak decode, which flatters the silicon; reviews print single-stream, which flatters the reviewer’s turnaround time. The ladder takes hours to measure properly and does not produce one heroic number, which is exactly why it is trustworthy and exactly why nobody leads with it.
A request is not an employee — real agents idle, burst, and paste enormous contexts at the worst moments, so a closed-loop ladder is a lower bound on chaos, not a forecast. But a lower bound beats a soloist’s headshot. If the swarm era is actually arriving, the spec that matters is agents-per-box at a latency floor, and the sooner we start asking for it, the sooner someone starts printing it.
Our measured ladders, with raw runs: the concurrency report · sustained load over three hours · the claim registry.