AGmind Systems Lab
Local LLMs on real hardware. Measured.
How fast a 35B-class model answers on a 128 GB box, what a pair of DGX Sparks really serves, which quant and backend to pick — measured on hardware we own, with raw runs published and failures kept in the denominator.
A device spec says how much memory a box has. A demo benchmark gives one number. Neither answers whether your exact model and runtime can be operated under your real load. AGmind Systems Lab tests the whole system and limits every conclusion to what was actually run.
01Evidence
Reviewed configurations
- PASS WITH LIMITS activeStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — document session cache
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Workload:
- doc-session-v1 @ 2026-08-03 — shared-prefix question pairs over 8k/32k-token documents, EN+RU
- PASS WITH LIMITS activeStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — three hours under sustained load
- System:
- 2× Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified — one pass per unit, side by side
- Workload:
- endurance-30m-v1 @ 2026-08-03 — human-task corpus cycled closed-loop for 180 minutes, checkpoints 30/60/180 min
- PARTIAL activeStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — everyday assistant
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified — single-user cells replicated on a second commercially identical unit
- Workload:
- interactive-assistant-v2 @ 2026-08-02 — 16 everyday tasks, EN+RU
- PASS WITH LIMITS activeStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — pasted-document retrieval
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Workload:
- long-context-v1 @ 2026-08-03 — needle ladder, nominal 2k–32k rungs (EN measured 1.8k–30.0k, RU 1.4k–21.1k), with controls
- PASS WITH LIMITS activeStrix Halo × llama.cpp Vulkan × Qwen3.6-35B-A3B — strict JSON automation
- System:
- Beelink GTR9 Pro — Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB unified
- Workload:
- structured-agent-v1 @ 2026-08-03 — 16 deterministic JSON tasks, EN+RU
Latest reports
All configurations →- Beelink GTR9 Pro for local LLMs: what two units measured over a month of serving
Reviews of this box benchmark it for a week. We bought two, froze the workloads, and have been serving from them since: the software stack that works, what the numbers look like, how the two units compare, and the parts we still cannot vouch for.
- vLLM at 256K context on DGX Spark: the configuration hunt, and the NVFP4 path that was broken in mainline
The 256K-context configuration hunt on a single GB10: which vLLM attention and MoE backends survived SM_121, why NVFP4 decoded slower than FP8 in the mainline build we pinned, the patched setup that held 260K — and the measured numbers, dated to the versions named inside.
- Local LLM homelab hardware, measured: what to buy for what workload
Every hardware guide for local AI is a spec sheet, a vendor reviewing itself, or an affiliate listicle recycling VRAM tiers. This one starts from the workload, uses only numbers we measured on hardware we own, and marks unmeasured lanes as gaps.
- A Russian retrieval embedder in 24M parameters — and the one number that nearly ruined it
The story behind strizh-ru-retriever: layer-pruning RuModernBERT-small into a four-layer encoder, the distillation shortcut that produced a beautiful loss and dead retrieval, what the compression kept — and the leftover config value that silently truncated every input.
- A Russian RAG splitter that cuts documents by indices, not by rewriting the text
The story behind agmind-rag-splitter-ru: why a chunker that rewrites text is dangerous for RAG, how a LoRA fine-tune learned to return boundary indices as JSON, the training corpus published alongside it — and the honest gap between teacher agreement and retrieval quality.
- Strix Halo memory allocation: why ROCm sees a fraction of your 128 GB, and what to set instead
The classic first failure on a Ryzen AI Max+ 395 box: the model does not fit, or the runtime reports a few gigabytes of VRAM on a 128 GB machine. The BIOS split is usually the wrong lever. Here is what our serving node actually reports, and the setting that matters.
02Guides
Deployment guides
- Deploy a large MoE on DGX Spark with vLLM: single node and a two-node pair
The deployment order that avoids the silent hangs we hit on our own GB10 pair: weights and image pinned, head before peer, readiness by endpoint not by feel, and the kernel-selection check that separates the fast configuration from the one most people run.
2-3 h
- Deploy a local LLM on Strix Halo (Ryzen AI Max+ 395): llama.cpp from zero to first answer
The exact path we use on our own boxes: memory setup the BIOS won't tell you about, a pinned llama.cpp container on Vulkan, and the three checks that prove the server actually works — including the two failure modes that hide behind green statuses.
45 min
- Ollama on Strix Halo (Ryzen AI Max+ 395): install, GPU detection, first model
The default-path Ollama setup on a Strix Halo box, with the one log line that confuses everyone decoded: why the Vulkan path drops your iGPU, why ROCm picks it up anyway, and how to verify the model actually landed on the GPU instead of silently running on CPU.
20 min
03Integrity
Failures and corrections
Negative results are first-class output here: a paid test that refutes a claim is published (in independent mode) with the same rigor as a pass. The errata log records every correction to published results.
Read the errata policy →04Method
Methodology
Every qualification follows the same controlled process: an agreed question, frozen versions and workload, metrics and exclusion rules fixed before the runs (preregistered publicly from the flagship onward), preserved failures and invalid runs, a scoped conclusion, and a reproducible evidence bundle. Headline numbers carry an evidence level stating how much measurement stands behind them — single-run results are labeled as such, never presented as repeated; failed requests stay in the denominator.
Read the methodology →05Testbed
Lab hardware
The stands that actually exist in the lab, with their limits stated. No result generalizes beyond the tested unit and versions.
| Stand | Spec | Role | Limit | Status |
|---|---|---|---|---|
| 2× Beelink GTR9 Pro — AMD Strix Halo | Ryzen AI Max+ 395, 128 GB LPDDR5X-8000 unified, Radeon 8060S (gfx1151), dual 10GbE — two commercially identical units | Flagship qualification target: backend comparisons (Vulkan vs ROCm), unit-to-unit replication, multi-slot serving, RAG side-services | Two units do not represent the whole production batch | online |
| 2× NVIDIA DGX Spark — GB10 | GB10 Grace Blackwell, 20-core Arm, 128 GB unified, sm_121, ConnectX-7 — linked point-to-point over 200G RoCE | aarch64/sm_121 portability, multi-node topologies, long-context and speculative decoding studies | Expensive narrow testbed; results do not generalize to datacenter Blackwell | online |
| RTX 5090 workstation | Consumer Blackwell, 32 GB GDDR7, x86 host | CUDA control lane, fine-tuning/distillation, consumer-GPU baselines | One configuration; not an enterprise server | online |
| Apple M1 Max, 64 GB | Apple Silicon, 64 GB unified, macOS / MLX lane | Apple/MLX smoke tests and small cross-platform anchors | Not the current high-end Apple generation | online |
| MacBook Pro — Apple M4 Pro, 24 GB | Apple Silicon M4 Pro, 24 GB unified, macOS / MLX lane | Current-generation Apple anchor: laptop-class local inference, MLX vs llama.cpp Metal comparisons | 24 GB caps model size; thermals of a laptop, not a desktop | online |
| Laptop — Ryzen 9 9955HX + RTX 5070 Ti | Ryzen 9 9955HX (Zen 5, 16C), RTX 5070 Ti Laptop 12 GB GDDR7, 32 GB DDR5 | Discrete-GPU laptop lane: mobile Blackwell CUDA, VRAM-constrained inference and offload studies vs unified-memory machines | 12 GB VRAM forces offload for mid-size models; laptop thermals | online |
06Services
What AGmind qualifies
-
Vendor Claim Evidence Pack
Is one specific public claim about a system true under stated conditions?
Evidence-backed verdict with raw artifacts, limitations, and reproduction recipe.
-
Workload Capacity Qualification
What load does this exact system sustain under our workload and SLO?
Operating envelope: recommended range, pass-with-limits range, SLO-fail boundary, known failure modes.
-
Runtime/Model Portability Sprint
Can our exact runtime/model be made to work on Strix Halo, GB10, RTX or Apple Silicon — and what does it take?
Diagnosis with exact failure points, then (optionally) patches/flags/build recipes labeled as commissioned engineering.
-
Long-Context Reliability Qualification
How much context is actually useful on this system — not just how much fits in memory?
Useful-context verdict per depth with quality evidence, latency profile, and failure documentation.
-
Procurement Decision Pack
Which of 2–4 candidate systems should we buy for our workload?
Side-by-side evidence with explicit limits, not a universal score.
-
Air-Gapped Readiness Check
Will this stack install and operate in an isolated network segment?
Verified technical-control checklist with evidence per control.
Deliberately not offered yet
These become products only after paid demand proves they should exist:
- AGmind Reference Stack after demand gate
- Release Revalidation Channel after demand gate
- BenchOps after demand gate
07Workloads
Workload library
Qualification runs against fixed, versioned workloads — not ad-hoc prompts. All v1 workloads are drafts until their corpora, hashes and acceptance checks are frozen.
- interactive-assistant-v1 released
Is single-user streaming chat responsive on this system?
- interactive-assistant-v2 released
Is single-user streaming chat responsive on this system for everyday human tasks?
- team-serving-v1 Draft
What request rate does the system sustain within a stated SLO?
- long-context-v1 released
How deep is the useful context on this system, not the configurable one?
- structured-agent-v1 released
Does strict JSON output stay reliable for everyday automation tasks on this system?
- endurance-30m-v1 released
Does this system degrade under sustained load — and when?
- rag-pipeline-ru-v0 Internal
Where does a Russian-language RAG pipeline lose quality, component by component?
- doc-session-v1 released
Does the second question over the same pasted document pay the full prefill again on this system?
08Scope
Commission a qualification
You buy a controlled process, not a positive result. Fees never depend on the verdict, funding is disclosed, and independent tests are separated from commissioned engineering.