The question this card answers
Mini-PCs throttle. Does this one, under a realistic sustained assistant load — and does the answer hold on more than one unit?
Verdict: PASS WITH LIMITS. Over three continuous hours at closed-loop concurrency 4 neither unit degraded measurably or dropped a single request. The evidence level is honest about depth: one 3-hour pass per unit.
Results
Completion over the full pass, both units pooled: 100.0% of requestssingle_run. The decode pace the units sustained throughout: 30.0ms/tokensingle_run. The number this workload exists to catch — the drift of that pace between the first five minutes and minutes 175–180: 1.6%single_run — within noise, at every checkpoint (30, 60, 180 min) on both units.
The thermal picture behind it, from a die-temperature sidecar log (documented per run, ambient not instrumented): unit A settled around the low seventies Celsius, unit B ran roughly six degrees hotter at an identical request pace — unit-to-unit variance in cooling that, at least within this envelope, did not translate into a performance difference.
The limits in “pass with limits”
- One 3-hour pass per unit. The pattern reproduced on both commercially
identical units, but within-unit repeats are pending — hence
lab_single_run. - This envelope only: closed-loop concurrency 4, reasoning off, this corpus, a decode pace of 30.0ms/tokensingle_run. A heavier prefill-dominated load could behave differently.
- Not a thermals verdict. Die temperature was logged, ambient was not instrumented; no claim about hotter rooms, dust, or seasons.
- The word-budget verbosity gate keeps failing on this corpus, as on the everyday assistant card — sustained load did not change output discipline either way.
Revalidation triggers
Runtime image change, model artifact change, workload revision change, driver/kernel change, chassis/cooling change on either unit. Harness and corpus: agmind-bench.