M1 Pro (16-core GPU) LLM benchmark
The fastest LLM measured on the M1 Pro (16-core GPU) is llama3.2 at 103.1 decode tok/s via ollama (signed run). Across 4 reproducible runs on 1 model, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.
Fastest known config on M1 Pro (16-core GPU)
103.1 decode tok/s
llama3.2 via ollama (Q8_0). see full run
llama3.2
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.32.14 | Q8_0 | 103.1tok/s | 8.77tok/s | 14,247ms | r_w-p7v--eync |
| chat-long | ollama@0.32.14 | Q8_0 | 81.35tok/s | 1,246.5tok/s | 2,529ms | r_w-p7v--eync |
| concurrent-decode | ollama@0.32.14 | Q8_0 | 97.26tok/s | no data | no data | r_w-p7v--eync |
| agent-trace | ollama@0.32.14 | Q8_0 | 93.14tok/s | 2,975.3tok/s | 516ms | r_w-p7v--eync |
Models measured on M1 Pro (16-core GPU)
Common questions about M1 Pro (16-core GPU)
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.