Claim

Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan) — time to first answer token, 95th percentile: 302 ms

active lab_unit_replicated strix.qwen36.interactive2.c1.ttfa-p95-nothink

302ms

The measurement

On a Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified running llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14) with ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M), the measured time to first answer token, 95th percentile was 302 ms (95th percentile over valid requests, interpolated). The measurement was taken under the frozen interactive-assistant-v2@2026-08-02 workload; evidence level lab_unit_replicated, 6 valid runs across 2 physical units. The value is re-derived from the raw run records on every CI build.

Statement

95th-percentile client-side time to the first token of the answer with the reasoning block disabled, human-task corpus — the slow tail of the same requests behind the median claim.

Evidence

System
Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified
Runtime
llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14)
Model artifact
ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M)
Scope
interactive-assistant-v2@2026-08-02
Aggregation
95th percentile over valid requests, interpolated
Evidence level
lab_unit_replicated
Status
active
Published
September 2, 2026
Limitations
Three repeated runs across two commercially identical units. Each run has 16 corpus items, so even pooled across the listed runs the 95th percentile rests on a handful of the slowest requests — a coarse tail statistic, not a tail-latency model. Failed requests carry no answer-token time and are outside this statistic; the completion-share claim on the same runs accounts for them.
Runs

Derivation

The value is re-derived from the raw run records by this query on every CI build — a published number cannot silently drift from its evidence.

catalog/claims/sql/ttfa-p95-nothink.sql

Cited on

Pages on this site that render this number:

Cite this claim

AGmind Systems Lab. Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan): time to first answer token, 95th percentile — 302 ms (95th percentile over valid requests, interpolated; 6 runs on 2 units; evidence level lab_unit_replicated; workload interactive-assistant-v2@2026-08-02). Claim strix.qwen36.interactive2.c1.ttfa-p95-nothink. https://agmind.ai/claims/strix.qwen36.interactive2.c1.ttfa-p95-nothink/
@misc{agmind_strix_qwen36_interactive2_c1_ttfa_p95_nothink,
  author       = {{AGmind Systems Lab}},
  title        = {95th-percentile client-side time to the first token of the answer with the reasoning block disabled, human-task corpus — the slow tail of the same requests behind the median claim.},
  howpublished = {\url{https://agmind.ai/claims/strix.qwen36.interactive2.c1.ttfa-p95-nothink/}},
  note         = {Claim strix.qwen36.interactive2.c1.ttfa-p95-nothink: 302 ms; evidence level lab\_unit\_replicated; scope interactive-assistant-v2@2026-08-02},
  year         = {2026}
}

Machine-readable: /claims/strix.qwen36.interactive2.c1.ttfa-p95-nothink.json · BibTeX · CSL-JSON · full registry · changes feed

Own comparable hardware? This claim can be reproduced: /reproduce/

Corrections to published results are logged publicly: errata

← All claims