Skip to content
llm-speed
Leaderboard/hardware/m5-max-40-core-gpu

M5 Max (40-core GPU) LLM benchmark

The fastest LLM measured on the M5 Max (40-core GPU) is qwen3.6-agent-q6k-ctx128k at 88.3 decode tok/s via ollama (signed run). Across 1 reproducible run on 1 model, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.

Fastest known config on M5 Max (40-core GPU)

88.3 decode tok/s

qwen3.6-agent-q6k-ctx128k via ollama (Q6_K). see full run

qwen3.6-agent-q6k-ctx128k

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortollama@0.32.1Q6_K88.29tok/s16.20tok/s6,915msr_udwgba7udqu

Models measured on M5 Max (40-core GPU)

Common questions about M5 Max (40-core GPU)

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the M5 Max (40-core GPU) FAQ →