Column

Agentic coding is a prefill problem, and everyone is shopping for decode

Ask how many tokens per second a local coding agent needs and every answer quotes decode speed. But an agent's turn is a huge prompt and a short reply — the wait lives in prefill, and the hardware conversation is optimizing the wrong half of the request.

opinion September 1, 2026

This is an opinion column, not a lab result. Numbers here come from the author’s working journal and quoted external sources (marked in the text) — never from the claim registry. Measured, evidence-backed results live at /claims/.

Every week someone asks how many tokens per second they need for a local coding agent, and every week the answers quote decode speed. It is the wrong half of the request. Look at what an agent actually sends: a system prompt, a tool manifest, file contents, build output, conversation history — tens of thousands of tokens of context — followed by a reply that is often one tool call long. The generation is a rounding error. The reading is the work.

That reading is prefill, and prefill behaves nothing like decode. Decode is a steady drip you can watch; prefill is a wall you wait behind before the first token appears. On our 128 GB box, a question over a 32k-token document costs tens of seconds of prefill before one visible character — the exact figure, measured, is in the context comparison — while the decode that follows runs faster than anyone reads. A coding agent pays something like that wall on every turn where its context changed. Buy hardware by the decode number and you have optimized the part of the turn your agent spends the least time in.

It gets worse before it gets better. Reasoning models stack a second wall behind the first: after the prefill, the model thinks, and the thinking stream is not addressed to you. We measured a two-hundred-millisecond first word turn into a twenty-second one with reasoning on — the thinking report has the pair — and an agent that leaves reasoning on by default pays both walls, every turn, mostly for tasks with checkable answers where our gates found reasoning bought nothing.

Now the better part, because there is one, and it is the most underappreciated lever in local serving: the prefix cache. An agent’s context is not arbitrary — it is the same system prompt, the same tool manifest, the same file headers, turn after turn, with changes at the tail. A server that caches the shared prefix pays the wall once and rides for the rest of the session. We measured the difference between paying and riding — the second question comes back almost immediately — and the condition it demands is brutal in its simplicity: the prefix must be byte-identical. One reshuffled retrieval chunk, one timestamp in the system prompt, and every turn pays full price while your dashboard shows a healthy cache.

Which reframes the whole hardware question. For agentic coding, the specs that matter are prefill throughput, and — more than anything — whether your agent framework and server keep the prefix stable enough to cache. A mid-tier box with a disciplined context layout will out-feel a monster whose framework rewrites the prompt every turn. That is a software property wearing a hardware costume, which is why no GPU review will ever tell you about it.

So before shopping: measure one real agent turn on the box you have — first visible token, not first thinking token, per turn, with your actual context sizes. The method is written up. If the wait is the problem, check what your stack does to the prefix before you check prices on a bigger card. The wall is real, but half the people staring at it built it themselves.

← All columns