M5 Max (40-core GPU) LLM benchmark
The fastest LLM measured on the M5 Max (40-core GPU) is qwen3.6-agent-q6k-ctx128k at 88.3 decode tok/s via ollama (signed run). Across 1 reproducible run on 1 model, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M5 Max (40-core GPU)
88.3 decode tok/s
qwen3.6-agent-q6k-ctx128k via ollama (Q6_K). see full run
qwen3.6-agent-q6k-ctx128k
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.32.1 | Q6_K | 88.29tok/s | 16.20tok/s | 6,915ms | r_udwgba7udqu |
Models measured on M5 Max (40-core GPU)
Common questions about M5 Max (40-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.