# Deploy a local LLM on Strix Halo (Ryzen AI Max+ 395): llama.cpp from zero to first answer

> The exact path we use on our own boxes: memory setup the BIOS won't tell you about, a pinned llama.cpp container on Vulkan, and the three checks that prove the server actually works — including the two failure modes that hide behind green statuses.

- Published: 2026-09-01
- Platform: Strix Halo (Ryzen AI Max+ 395)
- Time to complete: 45 min
- Canonical: https://agmind.ai/guides/deploy-llm-strix-halo/

This is the deployment path we run on our own Ryzen AI Max+ 395 machines —
the same boxes behind [our published numbers](https://agmind.ai/reports/beelink-gtr9-pro-local-llm/).
No theory, no options catalog: one working path, with the traps marked where
we hit them. Performance expectations are deliberately absent — those live
in [the reports](https://agmind.ai/reports/), measured properly.

You need: a Strix Halo box (any vendor), Linux with a recent kernel, Docker,
and disk space for the model you pick.

## Step 1 — fix the memory ceiling first

On Linux your models will live in GTT, not in the BIOS-dedicated VRAM slice,
and the ceiling is a kernel parameter. Skipping this step is the first thing to rule out when a large model refuses to load on this hardware. Check the
current limit:

```
cat /sys/module/ttm/parameters/pages_limit
```

Multiply by 4 KiB to get bytes. If it is a fraction of your RAM, raise it via
the kernel command line (`ttm.pages_limit=...`) and reboot. The full
explanation of why the BIOS slider is the wrong lever — with readings from
our serving node — is in
[the memory allocation report](https://agmind.ai/reports/strix-halo-memory-allocation/).
Leave the BIOS slice small.

## Step 2 — get a model

Pick a GGUF artifact and note its exact revision. A 35B-class MoE at Q4_K_M
is the sweet spot we serve daily on this hardware — it leaves room for
context and a second model. Which quant to pick is
[a measured question](https://agmind.ai/reports/q8-0-vs-q4-k-m-strix-halo/); short version:
start with Q4_K_M, not the biggest file that fits.

```
mkdir -p /var/lib/llm/models
# download your chosen .gguf into it, and record WHERE it came from —
# repository and revision. "main" moves; your notes should not.
```

## Step 3 — run the server, pinned

The Vulkan build of llama.cpp's server runs on this iGPU out of the box:

```
docker run -d --name llm-server \
  --device /dev/dri --device /dev/kfd \
  -v /var/lib/llm/models:/models \
  -p 8080:8080 \
  ghcr.io/ggml-org/llama.cpp:server-vulkan \
  -m /models/YOUR-MODEL.gguf \
  -ngl 999 -fa on -c 32768 \
  --host 0.0.0.0 --port 8080 --jinja
```

Two habits that will save you later. First, after it works, re-pull by
digest (`ghcr.io/ggml-org/llama.cpp@sha256:...`) instead of the tag — tags
move, and when a future update changes your behavior you will want to know
exactly which build you were on. Second, if you change `--port`, know that
the image's built-in healthcheck still probes 8080: your container will show
`unhealthy` while serving perfectly. We wrote up
[that exact trap](https://agmind.ai/essays/dashboard-lies-both-directions/) after it bit us —
either keep the internal port at 8080 or override the healthcheck.

## Step 4 — ask the server what it actually runs

Do not trust your command line; trust the process:

```
curl -s localhost:8080/props | python3 -m json.tool | head -40
```

Read `n_ctx`, the model path, and the slot count. Defaults change between
llama.cpp versions, and the flag you did not set is still part of your
configuration — [we learned this the expensive way](https://agmind.ai/essays/luddites-were-right/).

## Step 5 — the smoke test that actually proves something

An HTTP 200 proves nothing yet. Send a real request and check three things:

```
curl -s localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Reply with one short sentence: what is inference?"}],
       "max_tokens":400,"temperature":0}' \
  | python3 -c "import json,sys; d=json.load(sys.stdin); m=d['choices'][0]['message']; \
    print('content:', repr((m.get('content') or '')[:120])); \
    print('finish:', d['choices'][0].get('finish_reason')); \
    print('tokens:', d.get('usage',{}).get('completion_tokens'))"
```

Check one: `content` is **not empty**. A reasoning model can burn the whole
token budget thinking and return an empty answer under a perfect HTTP 200 —
[we measured how often](https://agmind.ai/reports/answerless-http-200/), and if you see it,
raise `max_tokens` or disable thinking. Check two: `finish` is `stop`, not
`length`. Check three: count tokens from `usage.completion_tokens` if you
ever measure anything — never from the number of streaming events.

## Step 6 — decide about reasoning mode

If your model has a thinking mode, decide deliberately whether it stays on.
For everyday chat and structured automation on this hardware we measured
[what it costs and what it bought](https://agmind.ai/reports/should-you-disable-thinking-local-llm/) —
read that before shipping the default to your users.

## Where to go next

- **Backend choice**: Vulkan is our serving default on this hardware; the
  [ROCm comparison](https://agmind.ai/compare/llamacpp-vulkan-vs-rocm-strix-halo/) shows what
  each buys.
- **More users**: before promising the box to a team, read the
  [concurrency ladder](https://agmind.ai/reports/concurrency-capacity-strix-halo/).
- **Verify your own numbers**: the workloads behind our reports are
  [public](https://github.com/botAGI/agmind-bench) — run them against your
  fresh server and [compare against our claims](https://agmind.ai/claims/).

---

Machine-readable claim registry: https://agmind.ai/claims.json · llms.txt: https://agmind.ai/llms.txt
