Lab

About AGmind Systems Lab

An engineering lab that qualifies local AI systems under fixed, versioned workloads and publishes reproducible, evidence-backed results.

What the lab does

A device specification says how much memory and compute a machine has. A demo benchmark shows one isolated number. Neither answers the operational question: can this exact model and runtime be operated under a real workload on this exact machine. The practical risks live between the layers — a runtime formally supports the architecture but the exact build lacks the needed kernels; a model loads but long-context quality degrades; a configuration passes a health check but discovery, auth or restore does not work.

AGmind qualifies the system as a whole. A qualification fixes:

  • exact hardware and exact software fingerprint;
  • exact model artifact;
  • a frozen workload with functional and quality gates;
  • performance and reliability measurements;
  • an operating envelope and a deployment recipe;
  • an evidence bundle behind every published claim.

Every conclusion is limited to the tested workload, versions and quality gates. The method is documented on the methodology page; paid engagements are described under qualification.

What the lab is not

  • Not a SaaS chat product. AGmind does not host client inference and does not store client documents as a service.
  • Not GPU hosting. Lab hardware runs lab tests, not customer workloads.
  • Not a universal integrator. No “AI for every department” projects and no unscoped consulting.
  • Not paid positive reviews. Payment never depends on the verdict, and in independent mode negative results are published, not hidden.
  • Not formal certification. AGmind never brands results with certification wording.

How independent tests are separated from commissioned engineering, and how funding is disclosed, is defined on the independence page.

How the lab operates

The business loop is deliberately simple:

  1. own research on lab hardware;
  2. a public, evidence-backed report;
  3. a fixed-scope paid qualification for a specific system owner;
  4. a repeatable qualification process and a reference configuration;
  5. paid revalidation when a critical layer changes.

Funding is always disclosed. Negative results are first-class outcomes: they are paid for and published the same way as positive ones. Corrections are tracked on the errata page, and every headline number is generated from the claim registry described on the data page.

Lab testbed

Physical machines the lab owns and operates. Each entry states what the node is used for and what its results cannot be generalized to.

Node Specification Role Limits Status
2× Beelink GTR9 Pro — AMD Strix Halo Ryzen AI Max+ 395, 128 GB LPDDR5X-8000 unified, Radeon 8060S (gfx1151), dual 10GbE — two commercially identical units Flagship qualification target: backend comparisons (Vulkan vs ROCm), unit-to-unit replication, multi-slot serving, RAG side-services Two units do not represent the whole production batch online
2× NVIDIA DGX Spark — GB10 GB10 Grace Blackwell, 20-core Arm, 128 GB unified, sm_121, ConnectX-7 — linked point-to-point over 200G RoCE aarch64/sm_121 portability, multi-node topologies, long-context and speculative decoding studies Expensive narrow testbed; results do not generalize to datacenter Blackwell online
RTX 5090 workstation Consumer Blackwell, 32 GB GDDR7, x86 host CUDA control lane, fine-tuning/distillation, consumer-GPU baselines One configuration; not an enterprise server online
Apple M1 Max, 64 GB Apple Silicon, 64 GB unified, macOS / MLX lane Apple/MLX smoke tests and small cross-platform anchors Not the current high-end Apple generation online
MacBook Pro — Apple M4 Pro, 24 GB Apple Silicon M4 Pro, 24 GB unified, macOS / MLX lane Current-generation Apple anchor: laptop-class local inference, MLX vs llama.cpp Metal comparisons 24 GB caps model size; thermals of a laptop, not a desktop online
Laptop — Ryzen 9 9955HX + RTX 5070 Ti Ryzen 9 9955HX (Zen 5, 16C), RTX 5070 Ti Laptop 12 GB GDDR7, 32 GB DDR5 Discrete-GPU laptop lane: mobile Blackwell CUDA, VRAM-constrained inference and offload studies vs unified-memory machines 12 GB VRAM forces offload for mid-size models; laptop thermals online

Open artifacts

The lab publishes its working code and models in the open. Highlights (≈4.9k total model downloads on Hugging Face as of 2026-07-31):

Research & evidence repos

Deployment stacks

  • AGmind64 — self-hosted LLM/RAG stack for Strix Halo / x86_64
  • AGmind — the same for NVIDIA DGX Spark / GB10 (arm64)
  • AGmind-macos — one-command local AI for macOS (Metal)

Models on Hugging Face

Model quality figures are deliberately absent here: per the lab's own rules they appear only as registry-backed claims after methodology-v1 evaluation.

Where the lab publishes

Every surface below is the same lab. Evidence lands on GitHub, artifacts on Hugging Face, long-form write-ups on Habr; agmind.dev is the sister project shipping the open-source installers this lab qualifies against.

Contact

The fastest way to a useful answer is a concrete question: which system, which workload, which versions. Write on Telegram at @AGmind or start from a structured scope request.