# Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan) — time to first token, second question over an 8k document, cache on: 660 ms

On a Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified running llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14) with ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M), the measured time to first token, second question over an 8k document, cache on was 660 ms (median over valid requests for this item). The measurement was taken under the frozen doc-session-v1@2026-08-03 workload; evidence level lab_repeated, 3 valid runs across 1 physical unit. The value is re-derived from the raw run records on every CI build.

## Statement

> Median time to first token for the second question over the same 8k-token document with the server prompt cache enabled.

## Evidence

| Field | Value |
| --- | --- |
| Claim id | `strix.qwen36.docsession.c1.ttft-q2-8k-cache` |
| Value | 660 ms |
| Metric | time to first token, second question over an 8k document, cache on |
| Aggregation | median over valid requests for this item |
| System | Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified |
| Runtime | llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14) |
| Model artifact | ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M) |
| Workload scope | doc-session-v1@2026-08-03 |
| Evidence level | lab_repeated |
| Status | active |
| Runs | 3 across 1 unit(s) |
| Published | 2026-08-03 |
| Last changed | 2026-08-03 |

**Limitations.** Three repeated runs on one unit, reasoning disabled, single stream. Cache benefit requires a byte-identical document prefix in the same server session.

## Raw evidence

- Sealed run bundle `run-20260803-ds-cache-a`: https://github.com/botAGI/agmind-lab/tree/main/runs/run-20260803-ds-cache-a
- Sealed run bundle `run-20260803-ds-cache-a-r2`: https://github.com/botAGI/agmind-lab/tree/main/runs/run-20260803-ds-cache-a-r2
- Sealed run bundle `run-20260803-ds-cache-a-r3`: https://github.com/botAGI/agmind-lab/tree/main/runs/run-20260803-ds-cache-a-r3
- Derivation query: https://github.com/botAGI/agmind-lab/blob/main/catalog/claims/sql/ttft-ds-en-8k-q2.sql
- Machine-readable: https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache.json · full registry https://agmind.ai/claims.json

## Cited on

- https://agmind.ai/reports/llamacpp-prompt-cache-strix-halo/
- https://agmind.ai/tested-configurations/strix-halo-qwen36-llamacpp-vulkan-docsession/
- https://agmind.ai/compare/prompt-cache-on-vs-off-doc-session/
- https://agmind.ai/measured/qwen3-6-35b-a3b-q4-k-m-on-ryzen-ai-max-395/

## Cite this

AGmind Systems Lab. Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan): time to first token, second question over an 8k document, cache on — 660 ms (median over valid requests for this item; 3 runs on 1 unit; evidence level lab_repeated; workload doc-session-v1@2026-08-03). Claim strix.qwen36.docsession.c1.ttft-q2-8k-cache. https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache/

Registry changes feed: https://agmind.ai/claims/changes.xml · Corrections to published results: https://agmind.ai/errata/ · Data licensed CC BY 4.0.
Source page: https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache/
