# Prompt cache on vs off: the second question over the same document

Does the second question over the same pasted document pay the full prefill again on a local llama.cpp server?

With the cache on, the second question over a large document answers in under a second; with it off, the session pays the full half-minute prefill again. For paste-a-document workflows this flag is the difference between an interactive tool and an unusable one.

**What was held fixed.** Same unit, backend, artifact; question pairs share a byte-identical document prefix. The only change is the server prompt-cache flag.

| Metric | Cache on | Cache off |
| --- | --- | --- |
| TTFT, 2nd question over the same 32k document | 860 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-32k-cache/)) | 33728 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-32k-nocache/)) |
| TTFT, 2nd question over the same 8k document (cache on) vs 1st question over a fresh 32k document | 660 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache/)) | 33940 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q1-32k/)) |

## Limits

Byte-identical prefixes only — the cache buys nothing for a different or edited document. Single user, single stream; multi-user cache isolation was not measured.

## Related evidence

- https://agmind.ai/reports/llamacpp-prompt-cache-strix-halo/
- https://agmind.ai/workloads/doc-session-v1/
- https://agmind.ai/tested-configurations/strix-halo-qwen36-llamacpp-vulkan-docsession/

Every number above is re-derived from sealed run bundles in CI: https://agmind.ai/claims.json (CC BY 4.0).
Source page: https://agmind.ai/compare/prompt-cache-on-vs-off-doc-session/
