Evidence
Tested configurations
One card = one exact system, runtime, model artifact and workload revision, with a verdict scoped to precisely that combination. A card never claims support for other versions of the same device.
- PASS WITH LIMITS active lab_repeatedStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — document session cache
The second question over the same pasted document: with the server prompt cache off it pays the full half-minute prefill again; with the cache on it answers in under a second. One setting turns a document workflow from unusable into fluid.
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Runtime:
- llama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
- Model:
- Qwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
- Workload:
- doc-session-v1 @ 2026-08-03 — shared-prefix question pairs over 8k/32k-token documents, EN+RU
- PASS WITH LIMITS active lab_single_runStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — three hours under sustained load
A continuous 3-hour closed-loop pass at concurrency 4 on both units: all requests answered, decode pace held between the first and the last five minutes, no thermal collapse — with the two units running at visibly different die temperatures.
- System:
- 2× Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified — one pass per unit, side by side
- Runtime:
- llama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
- Model:
- Qwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
- Workload:
- endurance-30m-v1 @ 2026-08-03 — human-task corpus cycled closed-loop for 180 minutes, checkpoints 30/60/180 min
- PARTIAL active lab_repeatedStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — everyday assistant
Single-user everyday chat on a Beelink GTR9 Pro: instant answers with reasoning disabled, a broken configuration at the default reasoning budget, and a verbosity limit that never fully clears.
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified — single-user cells replicated on a second commercially identical unit
- Runtime:
- llama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
- Model:
- Qwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
- Workload:
- interactive-assistant-v2 @ 2026-08-02 — 16 everyday tasks, EN+RU
- PASS WITH LIMITS active lab_repeatedStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — pasted-document retrieval
Needle retrieval across a context ladder up to 30k measured tokens with unanswerable controls: retrieval and honesty both held everywhere measured — while time to the first token grew linearly into tens of seconds.
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Runtime:
- llama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
- Model:
- Qwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
- Workload:
- long-context-v1 @ 2026-08-03 — needle ladder, nominal 2k–32k rungs (EN measured 1.8k–30.0k, RU 1.4k–21.1k), with controls
- PASS WITH LIMITS active lab_repeatedStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — strict JSON automation
Deterministic extraction, classification and structured generation in strict JSON: a perfect task-success rate in both reasoning modes on this corpus — with a seventeenfold end-to-end time difference between them.
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Runtime:
- llama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
- Model:
- Qwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
- Workload:
- structured-agent-v1 @ 2026-08-03 — 16 deterministic JSON tasks, EN+RU