Tested configuration

Strix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — three hours under sustained load

A continuous 3-hour closed-loop pass at concurrency 4 on both units: all requests answered, decode pace held between the first and the last five minutes, no thermal collapse — with the two units running at visibly different die temperatures.

PASS WITH LIMITS active lab_single_run

Exact configuration fingerprint

System2× Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified — one pass per unit, side by side
Runtimellama.cpp server-vulkan-b9049, pinned by image digest sha256:e359c012…
ModelQwen3.6-35B-A3B Q4_K_M (ggml-org, pinned revision, sha256:671e47e0…)
Workloadendurance-30m-v1 @ 2026-08-03 — human-task corpus cycled closed-loop for 180 minutes, checkpoints 30/60/180 min
Modeinternal research
FundingSelf-funded internal research

This verdict applies only to the exact configuration and workload revision above. It does not establish support for other firmware, driver, runtime or model versions of the same device.

The question this card answers

Mini-PCs throttle. Does this one, under a realistic sustained assistant load — and does the answer hold on more than one unit?

Verdict: PASS WITH LIMITS. Over three continuous hours at closed-loop concurrency 4 neither unit degraded measurably or dropped a single request. The evidence level is honest about depth: one 3-hour pass per unit.

Results

Completion over the full pass, both units pooled: 100.0% of requestssingle_run. The decode pace the units sustained throughout: 30.0ms/tokensingle_run. The number this workload exists to catch — the drift of that pace between the first five minutes and minutes 175–180: 1.6%single_run — within noise, at every checkpoint (30, 60, 180 min) on both units.

The thermal picture behind it, from a die-temperature sidecar log (documented per run, ambient not instrumented): unit A settled around the low seventies Celsius, unit B ran roughly six degrees hotter at an identical request pace — unit-to-unit variance in cooling that, at least within this envelope, did not translate into a performance difference.

The limits in “pass with limits”

Revalidation triggers

Runtime image change, model artifact change, workload revision change, driver/kernel change, chassis/cooling change on either unit. Harness and corpus: agmind-bench.

← Tested configurations