Claim

Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan) — share of output characters spent on reasoning: 86.5 %

active lab_unit_replicated strix.qwen36.interactive2.c1.reasoning-share-4k

86.5%

The measurement

On a Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified running llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14) with ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M), the measured share of output characters spent on reasoning was 86.5 % (median over valid requests of reasoning characters as a share of all output characters). The measurement was taken under the frozen interactive-assistant-v2@2026-08-02 workload; evidence level lab_unit_replicated, 6 valid runs across 2 physical units. The value is re-derived from the raw run records on every CI build.

Statement

Median share of the generated output that the model spent on the reasoning block rather than on the visible answer, counted in characters, with reasoning enabled at a 4096-token budget, human-task corpus.

Evidence

System
Beelink GTR9 Pro — AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X-8000 unified
Runtime
llama.cpp (server, Vulkan backend), b9049 (server_fingerprint b9049-2496f9c14)
Model artifact
ggml-org/Qwen3.6-35B-A3B-GGUF @ baec3ebee244 (Q4_K_M)
Scope
interactive-assistant-v2@2026-08-02
Aggregation
median over valid requests of reasoning characters as a share of all output characters
Evidence level
lab_unit_replicated
Status
active
Published
September 2, 2026
Limitations
Three repeated runs across two commercially identical units — diagnostic depth, not a cross-unit qualification. Counted in characters, not tokens: reasoning and answer text tokenize differently, so this is a proxy for how the token budget was spent, not the spend itself. Requests whose reasoning consumed the whole budget and left no answer are excluded here and counted by the answerless claim at the same 4096-token budget, which reads the same six runs — read the two together.
Runs

Derivation

The value is re-derived from the raw run records by this query on every CI build — a published number cannot silently drift from its evidence.

catalog/claims/sql/reasoning-share-4k.sql

Cited on

Pages on this site that render this number:

Cite this claim

AGmind Systems Lab. Qwen3.6-35B-A3B Q4_K_M on Ryzen AI Max+ 395 (llama.cpp Vulkan): share of output characters spent on reasoning — 86.5 % (median over valid requests of reasoning characters as a share of all output characters; 6 runs on 2 units; evidence level lab_unit_replicated; workload interactive-assistant-v2@2026-08-02). Claim strix.qwen36.interactive2.c1.reasoning-share-4k. https://agmind.ai/claims/strix.qwen36.interactive2.c1.reasoning-share-4k/
@misc{agmind_strix_qwen36_interactive2_c1_reasoning_share_4k,
  author       = {{AGmind Systems Lab}},
  title        = {Median share of the generated output that the model spent on the reasoning block rather than on the visible answer, counted in characters, with reasoning enabled at a 4096-token budget, human-task corpus.},
  howpublished = {\url{https://agmind.ai/claims/strix.qwen36.interactive2.c1.reasoning-share-4k/}},
  note         = {Claim strix.qwen36.interactive2.c1.reasoning-share-4k: 86.5 \%; evidence level lab\_unit\_replicated; scope interactive-assistant-v2@2026-08-02},
  year         = {2026}
}

Machine-readable: /claims/strix.qwen36.interactive2.c1.reasoning-share-4k.json · BibTeX · CSL-JSON · full registry · changes feed

Own comparable hardware? This claim can be reproduced: /reproduce/

Corrections to published results are logged publicly: errata

← All claims