# Thinking vs no-thinking: what reasoning costs in a chat

Should you leave the reasoning block on for everyday assistant tasks on a local Qwen3.6-35B-A3B?

For these everyday tasks reasoning multiplied the wait by two orders of magnitude and improved nothing the gates can measure: task success is identical with and without it. The user-visible difference is an instant assistant versus a twenty-second pause before the first answer token.

**What was held fixed.** Same unit, backend, artifact and corpora; the only change is the reasoning toggle (chat-template switch, 4096-token budget when on).

| Metric | Reasoning off | Reasoning on (4096) |
| --- | --- | --- |
| Time to first ANSWER token, median (chat) | 210 ms ([lab_unit_replicated](https://agmind.ai/claims/strix.qwen36.interactive2.c1.ttfa-nothink/)) | 20191 ms ([lab_unit_replicated](https://agmind.ai/claims/strix.qwen36.interactive2.c1.ttfa-thinking/)) |
| End-to-end time to complete JSON answer, median | 844 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.e2e-nothink/)) | 14677 ms ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.e2e-think4k/)) |
| Strict-JSON task success | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.task-success-nothink/)) | 100.0 % of requests ([lab_repeated](https://agmind.ai/claims/strix.qwen36.structured.c1.task-success-think4k/)) |

## Limits

Everyday-task and strict-JSON corpora only — reasoning may pay for itself on genuinely hard problems (math, multi-step planning) these workloads do not contain. The cost side is general; the benefit side is scoped.

## Related evidence

- https://agmind.ai/reports/ttft-thinking-model-strix-halo/
- https://agmind.ai/workloads/structured-agent-v1/

Every number above is re-derived from sealed run bundles in CI: https://agmind.ai/claims.json (CC BY 4.0).
Source page: https://agmind.ai/compare/thinking-vs-no-thinking-qwen36/
