# llama.cpp Vulkan vs ROCm on AMD Strix Halo

Which llama.cpp backend should you run on a Ryzen AI Max+ 395 (gfx1151) box for single-user chat — Vulkan or ROCm?

Vulkan decodes faster on this configuration; time to the first answer token is effectively a tie. Both backends passed every strict-JSON task in the structured workload, so the choice is about decode pace and operational preference, not quality.

**What was held fixed.** Same physical unit (Beelink GTR9 Pro, 128 GB unified), same Qwen3.6-35B-A3B Q4_K_M artifact by hash, same frozen human-task corpus, single stream, reasoning off. Only the backend container differs.

| Metric | Vulkan | ROCm |
| --- | --- | --- |
| Inter-token latency, median (decode pace) | 15.9 ms/token ([lab_repeated](https://agmind.ai/claims/strix.qwen36q4.vulkan.c1.itl-nothink/)) | 18.7 ms/token ([lab_repeated](https://agmind.ai/claims/strix.qwen36q4.rocm.c1.itl-nothink/)) |
| Time to first answer token, median | 210 ms ([lab_unit_replicated](https://agmind.ai/claims/strix.qwen36.interactive2.c1.ttfa-nothink/)) | 215 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36q4.rocm.c1.ttfa-nothink/)) |
| Strict-JSON task success (structured workload) | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.task-success-nothink/)) | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36q4.rocm.c1.task-success/)) |

## Limits

One unit, one model artifact, one runtime build per backend, single stream. Prefill-heavy or multi-stream loads may rank the backends differently — measure your cell before committing.

## Related evidence

- https://agmind.ai/reports/model-quant-backend-strix-halo/
- https://agmind.ai/workloads/interactive-assistant-v2/

Every number above is re-derived from sealed run bundles in CI: https://agmind.ai/claims.json (CC BY 4.0).
Source page: https://agmind.ai/compare/llamacpp-vulkan-vs-rocm-strix-halo/
