M4 Max (40-core GPU) LLM benchmark
The fastest LLM measured on the M4 Max (40-core GPU) is qwen3-coder-bench-32k at 113.3 decode tok/s via ollama (signed run). Across 17 reproducible runs on 2 models, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M4 Max (40-core GPU)
113.3 decode tok/s
qwen3-coder-bench-32k via ollama (Q4_K_M). see full run
qwen3-coder-bench-32k
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.30.11 | Q4_K_M | 113.3tok/s | 84.12tok/s | 1,308ms | r_roktphpc--8 |
| chat-long | ollama@0.30.11 | Q4_K_M | 97.40tok/s | 1,342.8tok/s | 2,344ms | r_roktphpc--8 |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.7tok/s | no data | no data | r_roktphpc--8 |
| agent-trace | ollama@0.30.11 | Q4_K_M | 103.2tok/s | 3,371.6tok/s | 477ms | r_roktphpc--8 |
qwen3-coder
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.30.11 | Q4_K_M | 94.51tok/s | 65.06tok/s | 1,691ms | r_r0di2hkku1h |
| chat-long | ollama@0.30.11 | Q4_K_M | 97.35tok/s | 1,397.3tok/s | 2,252ms | r_r0di2hkku1h |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.9tok/s | no data | no data | r_r0di2hkku1h |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_r0di2hkku1h |
| chat-short | ollama@0.30.11 | Q4_K_M | 55.45tok/s | 675.2tok/s | 163ms | r_1o8q4lhgj88 |
| chat-long | ollama@0.30.11 | Q4_K_M | 57.86tok/s | 28,422.4tok/s | 111ms | r_1o8q4lhgj88 |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 67.01tok/s | no data | no data | r_1o8q4lhgj88 |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_1o8q4lhgj88 |
| chat-short | ollama@0.30.11 | Q4_K_M | 92.70tok/s | 719.0tok/s | 153ms | r_40h1dznuznk |
| chat-long | ollama@0.30.11 | Q4_K_M | 94.61tok/s | 1,396.5tok/s | 2,253ms | r_40h1dznuznk |
| concurrent-decode | ollama@0.30.11 | Q4_K_M | 109.1tok/s | no data | no data | r_40h1dznuznk |
| agent-trace | ollama@0.30.11 | Q4_K_M | no data | no data | no data | r_40h1dznuznk |
| chat-short | ollama@0.30.11 | Q4_K_M | 93.17tok/s | 13.30tok/s | 8,270ms | r_txvkjgoyyxs |
Community folklore on M4 Max (40-core GPU)
72 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 75%
22.06tok/s — Phi4 on M4 Max via ollama fp16
our signed data: M4 Max · Phi4
“y **Results (All models downloaded from Ollama)** **gemma3:27b** |Quantization|Load Duration|Inference Speed| |:-|:-|:-| |q4|52.482042ms|22.06 tokens/s| |fp16|56.4445ms|6.99 tokens/s| **gemma3:12b** |Quantization|Load Duration|Inference Speed| |:-|:-|:-| |q4|56.818334ms|43.8…”
- communityconfidence 70%
25.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …”
- communityconfidence 70%
17.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s”
- communityconfidence 70%
25.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …”
- communityconfidence 70%
17.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s”
- communityconfidence 70%
25.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …”
- communityconfidence 70%
17.00tok/s — GPT-OSS 120B on M4 Max via lm-studio
our signed data: M4 Max · GPT-OSS 120B
“Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s”
- communityconfidence 60%
101.9tok/s — Qwen2.5-7B on M4 Max via mlx
our signed data: M4 Max · Qwen2.5-7B
“GPU) | | --------------------------- | -------------------------------- | ------------------------------ | | Qwen2.5-7B-Instruct (4bit) | 101.87 tokens/s | 38.99 tokens/s | | Qwen2.5-14B-Instruct (4bit) | 52.22 tokens/s | 18.88 …”
- communityconfidence 60%
8.76tok/s — Qwen2.5 on M4 Max via lm-studio
our signed data: M4 Max · Qwen2.5
“okens/s | | Qwen2.5:32B (4bit) | 19.35 tokens/s | 6.95 tokens/s | | Qwen2.5:72B (4bit) | 8.76 tokens/s | Didn't Test | #### LM Studio | MLX models | M4 Max (128 GB RAM, 40-…”
- communityconfidence 60%
6.95tok/s — Qwen2.5 on M4 Max via lm-studio
our signed data: M4 Max · Qwen2.5
“:14B (4bit) | 38.23 tokens/s | 14.66 tokens/s | | Qwen2.5:32B (4bit) | 19.35 tokens/s | 6.95 tokens/s | | Qwen2.5:72B (4bit) | 8.76 tokens/s | Didn't Test | #### LM Studio …”
Models measured on M4 Max (40-core GPU)
Common questions about M4 Max (40-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.