M3 Ultra (60-core GPU) LLM benchmark
The fastest LLM measured on the M3 Ultra (60-core GPU) is Qwen3-0.6B-4bit at 429.2 decode tok/s via mlx (signed run). Across 44 reproducible runs on 26 models, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M3 Ultra (60-core GPU)
429.2 decode tok/s
Qwen3-0.6B-4bit via mlx. see full run
Qwen3-30B-A3B-Instruct-2507-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 112.5tok/s | 308.7tok/s | 356ms | r_o6wbmfegmde |
| chat-long | mlx@0.31.3 | - | 96.37tok/s | 1,885.2tok/s | 1,669ms | r_o6wbmfegmde |
| concurrent-decode | mlx@0.31.3 | - | 105.9tok/s | no data | no data | r_o6wbmfegmde |
| agent-trace | mlx@0.31.3 | - | 101.3tok/s | 1,978.0tok/s | 1,057ms | r_o6wbmfegmde |
DeepSeek-R1-0528-Qwen3-8B-MLX-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 116.8tok/s | 424.5tok/s | 247ms | r_yy_6jfx70jq |
| chat-long | mlx@0.31.3 | - | 100.5tok/s | 1,104.0tok/s | 2,846ms | r_yy_6jfx70jq |
| concurrent-decode | mlx@0.31.3 | - | 110.6tok/s | no data | no data | r_yy_6jfx70jq |
| agent-trace | mlx@0.31.3 | - | 107.2tok/s | 1,108.9tok/s | 1,884ms | r_yy_6jfx70jq |
Qwen3-4B-Instruct-2507-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 182.7tok/s | 636.6tok/s | 173ms | r_p2rpb_0iyjj |
| chat-long | mlx@0.31.3 | - | 147.0tok/s | 1,913.7tok/s | 1,644ms | r_p2rpb_0iyjj |
| concurrent-decode | mlx@0.31.3 | - | 168.4tok/s | no data | no data | r_p2rpb_0iyjj |
| agent-trace | mlx@0.31.3 | - | 160.8tok/s | 1,913.6tok/s | 1,091ms | r_p2rpb_0iyjj |
Qwen3-8B-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 120.5tok/s | 446.1tok/s | 247ms | r_b1awwhkoo2v |
| chat-long | mlx@0.31.3 | - | 103.3tok/s | 1,108.8tok/s | 2,838ms | r_b1awwhkoo2v |
| concurrent-decode | mlx@0.31.3 | - | 114.3tok/s | no data | no data | r_b1awwhkoo2v |
| agent-trace | mlx@0.31.3 | - | 109.4tok/s | 1,109.1tok/s | 1,885ms | r_b1awwhkoo2v |
Qwen3-0.6B-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 429.2tok/s | 1,185.2tok/s | 92.8ms | r_6rgao6ao735 |
| chat-long | mlx@0.31.3 | - | 334.3tok/s | 6,893.3tok/s | 457ms | r_6rgao6ao735 |
| concurrent-decode | mlx@0.31.3 | - | 368.7tok/s | no data | no data | r_6rgao6ao735 |
| agent-trace | mlx@0.31.3 | - | 368.5tok/s | 8,617.7tok/s | 233ms | r_6rgao6ao735 |
Qwen3-Next-80B-A3B-Instruct-MLX-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 80.34tok/s | 24.49tok/s | 4,493ms | r_1pl79r50ofy |
| chat-long | mlx@0.31.3 | 4bit | 77.63tok/s | 1,608.6tok/s | 1,956ms | r_1pl79r50ofy |
| concurrent-decode | mlx@0.31.3 | 4bit | 78.60tok/s | no data | no data | r_1pl79r50ofy |
| agent-trace | mlx@0.31.3 | 4bit | 78.41tok/s | 1,586.3tok/s | 1,318ms | r_1pl79r50ofy |
qwen2.5-72b-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 16.31tok/s | 23.25tok/s | 5,635ms | r_5c80gthqlh6 |
llama-3.3-70b-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 16.78tok/s | 25.09tok/s | 5,420ms | r_sx3a4y9n-m4 |
stable-code-instruct-3b-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 192.5tok/s | 560.7tok/s | 226ms | r_y2_5y8oo97d |
Yi-Coder-9B-Chat-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 103.5tok/s | 390.9tok/s | 307ms | r_3hvui9a1yuc |
starcoder2-15b-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 64.16tok/s | 220.5tok/s | 490ms | r_wsxml_39dh_ |
granite-8b-code-instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 112.3tok/s | 525.7tok/s | 221ms | r_bue3bee0gw7 |
Codestral-22B-v0.1-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 47.49tok/s | 205.7tok/s | 559ms | r_79dvtag5fd_ |
DeepSeek-Coder-V2-Lite-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 168.3tok/s | 291.5tok/s | 449ms | r_l_v1-zq_qaz |
Qwen2.5-Coder-32B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 34.48tok/s | 144.1tok/s | 909ms | r_721b4bls_oq |
Qwen2.5-Coder-14B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 70.51tok/s | 306.1tok/s | 428ms | r_l36cijqxq4t |
Qwen2.5-Coder-7B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 138.6tok/s | 539.6tok/s | 243ms | r_uoehjq0nvc0 |
gpt-oss-20b-MXFP4-Q4
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 152.7tok/s | 239.9tok/s | 692ms | r_3ijun8ltjnb |
Qwen3-Coder-30B-A3B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 112.2tok/s | 204.0tok/s | 539ms | r_fpsca03u2o_ |
Qwen3-32B-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 34.41tok/s | 95.12tok/s | 1,156ms | r_anmmc80-aoq |
Qwen2.5-32B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 34.60tok/s | 141.9tok/s | 923ms | r_njgxtgyym1e |
gemma-2-9b-it-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 89.45tok/s | 170.7tok/s | 691ms | r_iz137eqvuzy |
Qwen2.5-14B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 70.85tok/s | 301.7tok/s | 434ms | r_v4bq1sviz4o |
phi-4-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 74.37tok/s | 255.0tok/s | 443ms | r_sqzp-0rdez- |
Llama-3.1-8B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 130.2tok/s | 400.3tok/s | 340ms | r_v2pbc0rq2l4 |
Qwen2.5-7B-Instruct-4bit
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 139.6tok/s | 190.0tok/s | 689ms | r_5r6rhiynenc |
Models measured on M3 Ultra (60-core GPU)
- Qwen3-30B-A3B-Instruct-2507-4bit benchmarks
- DeepSeek-R1-0528-Qwen3-8B-MLX-4bit benchmarks
- Qwen3-4B-Instruct-2507-4bit benchmarks
- Qwen3-8B-4bit benchmarks
- Qwen3-0.6B-4bit benchmarks
- Qwen3-Next-80B-A3B-Instruct-MLX-4bit benchmarks
- qwen2.5-72b-Instruct-4bit benchmarks
- llama-3.3-70b-Instruct-4bit benchmarks
- stable-code-instruct-3b-4bit benchmarks
- Yi-Coder-9B-Chat-4bit benchmarks
- starcoder2-15b-4bit benchmarks
- granite-8b-code-instruct-4bit benchmarks
- Codestral-22B-v0.1-4bit benchmarks
- DeepSeek-Coder-V2-Lite-Instruct-4bit benchmarks
- Qwen2.5-Coder-32B-Instruct-4bit benchmarks
- Qwen2.5-Coder-14B-Instruct-4bit benchmarks
- Qwen2.5-Coder-7B-Instruct-4bit benchmarks
- gpt-oss-20b-MXFP4-Q4 benchmarks
- Qwen3-Coder-30B-A3B-Instruct-4bit benchmarks
- Qwen3-32B-4bit benchmarks
- Qwen2.5-32B-Instruct-4bit benchmarks
- gemma-2-9b-it-4bit benchmarks
- Qwen2.5-14B-Instruct-4bit benchmarks
- phi-4-4bit benchmarks
- Llama-3.1-8B-Instruct-4bit benchmarks
- Qwen2.5-7B-Instruct-4bit benchmarks
Common questions about M3 Ultra (60-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.