# Q4_K_M vs Q8_0: does the bigger quant buy anything?

Is it worth spending the extra memory on the Q8_0 artifact of Qwen3.6-35B-A3B for chat and JSON automation on Strix Halo?

On the workloads these gates can see, the bigger artifact bought nothing: task success is identical, and Q8_0 decodes slower. If your workload resembles these — everyday chat and strict-JSON extraction — the memory is better spent on context.

**What was held fixed.** Same unit, same Vulkan backend build, same frozen corpora, single stream, reasoning off. Only the model artifact differs (both pinned by hash).

| Metric | Q4_K_M (20 GB) | Q8_0 (35 GB) |
| --- | --- | --- |
| Inter-token latency, median (decode pace) | 15.9 ms/token ([lab_repeated](https://agmind.ai/claims/strix.qwen36q4.vulkan.c1.itl-nothink/)) | 18.7 ms/token ([lab_repeated](https://agmind.ai/claims/strix.qwen36q8.vulkan.c1.itl-nothink/)) |
| Strict-JSON task success | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.task-success-nothink/)) | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36q8.vulkan.c1.task-success/)) |

## Limits

Quality gates here are format/JSON/repetition classes — a quant difference could surface on harder reasoning or multilingual generation these workloads do not probe. Negative result, scoped to what was measured.

## Related evidence

- https://agmind.ai/reports/q8-0-vs-q4-k-m-strix-halo/
- https://agmind.ai/reports/model-quant-backend-strix-halo/

Every number above is re-derived from sealed run bundles in CI: https://agmind.ai/claims.json (CC BY 4.0).
Source page: https://agmind.ai/compare/q4km-vs-q80-qwen36-strix-halo/
