AGmind Systems Lab

Local LLMs on real hardware. Measured.

How fast a 35B-class model answers on a 128 GB box, what a pair of DGX Sparks really serves, which quant and backend to pick — measured on hardware we own, with raw runs published and failures kept in the denominator.

A device spec says how much memory a box has. A demo benchmark gives one number. Neither answers whether your exact model and runtime can be operated under your real load. AGmind Systems Lab tests the whole system and limits every conclusion to what was actually run.

Reviewed configurations

All configurations →

Deployment guides

All guides →

Failures and corrections

Negative results are first-class output here: a paid test that refutes a claim is published (in independent mode) with the same rigor as a pass. The errata log records every correction to published results.

Read the errata policy →

Methodology

Every qualification follows the same controlled process: an agreed question, frozen versions and workload, metrics and exclusion rules fixed before the runs (preregistered publicly from the flagship onward), preserved failures and invalid runs, a scoped conclusion, and a reproducible evidence bundle. Headline numbers carry an evidence level stating how much measurement stands behind them — single-run results are labeled as such, never presented as repeated; failed requests stay in the denominator.

Read the methodology →

Lab hardware

The stands that actually exist in the lab, with their limits stated. No result generalizes beyond the tested unit and versions.

Stand Spec Role Limit Status
2× Beelink GTR9 Pro — AMD Strix Halo Ryzen AI Max+ 395, 128 GB LPDDR5X-8000 unified, Radeon 8060S (gfx1151), dual 10GbE — two commercially identical units Flagship qualification target: backend comparisons (Vulkan vs ROCm), unit-to-unit replication, multi-slot serving, RAG side-services Two units do not represent the whole production batch online
2× NVIDIA DGX Spark — GB10 GB10 Grace Blackwell, 20-core Arm, 128 GB unified, sm_121, ConnectX-7 — linked point-to-point over 200G RoCE aarch64/sm_121 portability, multi-node topologies, long-context and speculative decoding studies Expensive narrow testbed; results do not generalize to datacenter Blackwell online
RTX 5090 workstation Consumer Blackwell, 32 GB GDDR7, x86 host CUDA control lane, fine-tuning/distillation, consumer-GPU baselines One configuration; not an enterprise server online
Apple M1 Max, 64 GB Apple Silicon, 64 GB unified, macOS / MLX lane Apple/MLX smoke tests and small cross-platform anchors Not the current high-end Apple generation online
MacBook Pro — Apple M4 Pro, 24 GB Apple Silicon M4 Pro, 24 GB unified, macOS / MLX lane Current-generation Apple anchor: laptop-class local inference, MLX vs llama.cpp Metal comparisons 24 GB caps model size; thermals of a laptop, not a desktop online
Laptop — Ryzen 9 9955HX + RTX 5070 Ti Ryzen 9 9955HX (Zen 5, 16C), RTX 5070 Ti Laptop 12 GB GDDR7, 32 GB DDR5 Discrete-GPU laptop lane: mobile Blackwell CUDA, VRAM-constrained inference and offload studies vs unified-memory machines 12 GB VRAM forces offload for mid-size models; laptop thermals online

What AGmind qualifies

Deliberately not offered yet

These become products only after paid demand proves they should exist:

  • AGmind Reference Stack after demand gate
  • Release Revalidation Channel after demand gate
  • BenchOps after demand gate

Workload library

Qualification runs against fixed, versioned workloads — not ad-hoc prompts. All v1 workloads are drafts until their corpora, hashes and acceptance checks are frozen.

Commission a qualification

You buy a controlled process, not a positive result. Fees never depend on the verdict, funding is disclosed, and independent tests are separated from commissioned engineering.