M4 Max (40-core GPU) LLM benchmark
The fastest LLM measured on the M4 Max (40-core GPU) is qwen3-coder-bench-32k at 113.3 decode tok/s via ollama (signed run). Across 17 reproducible runs on 2 models, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M4 Max (40-core GPU)
113.3 decode tok/s
qwen3-coder-bench-32k via ollama (Q4_K_M). see full run
qwen3-coder-bench-32k
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.30.11 | Q4_K_M | 113.3tok/s | 84.12tok/s | 1,308ms | r_roktphpc--8 |
| chat-long | ollama@0.30.11 | Q4_K_M | 97.40tok/s | 1,342.8tok/s | 2,344ms | r_roktphpc--8 |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.7tok/s | no data | no data | r_roktphpc--8 |
| agent-trace | ollama@0.30.11 | Q4_K_M | 103.2tok/s | 3,371.6tok/s | 477ms | r_roktphpc--8 |
qwen3-coder
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.30.11 | Q4_K_M | 94.51tok/s | 65.06tok/s | 1,691ms | r_r0di2hkku1h |
| chat-long | ollama@0.30.11 | Q4_K_M | 97.35tok/s | 1,397.3tok/s | 2,252ms | r_r0di2hkku1h |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.9tok/s | no data | no data | r_r0di2hkku1h |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_r0di2hkku1h |
| chat-short | ollama@0.30.11 | Q4_K_M | 55.45tok/s | 675.2tok/s | 163ms | r_1o8q4lhgj88 |
| chat-long | ollama@0.30.11 | Q4_K_M | 57.86tok/s | 28,422.4tok/s | 111ms | r_1o8q4lhgj88 |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 67.01tok/s | no data | no data | r_1o8q4lhgj88 |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_1o8q4lhgj88 |
| chat-short | ollama@0.30.11 | Q4_K_M | 92.70tok/s | 719.0tok/s | 153ms | r_40h1dznuznk |
| chat-long | ollama@0.30.11 | Q4_K_M | 94.61tok/s | 1,396.5tok/s | 2,253ms | r_40h1dznuznk |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.1tok/s | no data | no data | r_40h1dznuznk |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_40h1dznuznk |
| chat-short | ollama@0.30.11 | Q4_K_M | 93.17tok/s | 13.30tok/s | 8,270ms | r_txvkjgoyyxs |
Models measured on M4 Max (40-core GPU)
Common questions about M4 Max (40-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.