Skip to content
llm-speed
Leaderboard/hardware/m1-pro-16-core-gpu

M1 Pro (16-core GPU) LLM benchmark

The fastest LLM measured on the M1 Pro (16-core GPU) is llama3.2 at 103.1 decode tok/s via ollama (signed run). Across 4 reproducible runs on 1 model, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.

Fastest known config on M1 Pro (16-core GPU)

103.1 decode tok/s

llama3.2 via ollama (Q8_0). see full run

llama3.2

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortollama@0.32.14Q8_0103.1tok/s8.77tok/s14,247msr_w-p7v--eync
chat-longollama@0.32.14Q8_081.35tok/s1,246.5tok/s2,529msr_w-p7v--eync
concurrent-decodeollama@0.32.14Q8_097.26tok/sno datano datar_w-p7v--eync
agent-traceollama@0.32.14Q8_093.14tok/s2,975.3tok/s516msr_w-p7v--eync

Models measured on M1 Pro (16-core GPU)

Common questions about M1 Pro (16-core GPU)

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the M1 Pro (16-core GPU) FAQ →