Claim

Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan) — time to first token, second question over an 8k document, cache on: 660 ms

active lab_repeated strix.qwen36.docsession.c1.ttft-q2-8k-cache

660ms

The measurement

On a Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified running llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14) with ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M), the measured time to first token, second question over an 8k document, cache on was 660 ms (median over valid requests for this item). The measurement was taken under the frozen doc-session-v1@2026-08-03 workload; evidence level lab_repeated, 3 valid runs across 1 physical unit. The value is re-derived from the raw run records on every CI build.

Statement

Median time to first token for the second question over the same 8k-token document with the server prompt cache enabled.

Evidence

System
Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified
Runtime
llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14)
Model artifact
ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M)
Scope
doc-session-v1@2026-08-03
Aggregation
median over valid requests for this item
Evidence level
lab_repeated
Status
active
Limitations
Three repeated runs on one unit, reasoning disabled, single stream. Cache benefit requires a byte-identical document prefix in the same server session.
Runs

Derivation

The value is re-derived from the raw run records by this query on every CI build — a published number cannot silently drift from its evidence.

catalog/claims/sql/ttft-ds-en-8k-q2.sql

Cite this claim

AGmind Systems Lab. "Median time to first token for the second question over the same 8k-token document with the server prompt cache enabled." Claim strix.qwen36.docsession.c1.ttft-q2-8k-cache (660 ms), evidence level lab_repeated, scope doc-session-v1@2026-08-03. https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache/
@misc{agmind_strix_qwen36_docsession_c1_ttft_q2_8k_cache,
  author       = {{AGmind Systems Lab}},
  title        = {Median time to first token for the second question over the same 8k-token document with the server prompt cache enabled.},
  howpublished = {\url{https://agmind.ai/claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache/}},
  note         = {Claim strix.qwen36.docsession.c1.ttft-q2-8k-cache: 660 ms; evidence level lab\_repeated; scope doc-session-v1@2026-08-03},
  year         = {2026}
}

Machine-readable: /claims/strix.qwen36.docsession.c1.ttft-q2-8k-cache.json · full registry

Own comparable hardware? This claim can be reproduced: /reproduce/

Corrections to published results are logged publicly: errata

← All claims