M1 Max LLM benchmark
No M1 Max LLM benchmarks yet. Be the first to submit a signed, reproducible tok/s run from your own M1 Max.
Will it run local LLMs?
No signed M1 Max run yet, but its 64GB of unified memory tells you what fits. At 4-bit it comfortably runs models up to about 57B parameters with room for context, and up to roughly 91B with tighter quantization. Check a specific model with the VRAM-fit tool.
For measured decode speed on the nearest hardware we have benchmarked: M3 Ultra · M4 Max · M3 Pro. Run the suite on your M1 Max to fill this page in.
No M1 Max benchmarks yet.
Run on YOUR hardware to populate this page: pipx install llm-speed && llm-speed bench
$ pipx install llm-speed && llm-speed bench
Community folklore on M1 Max
13 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 75%
10.62tok/s — qwen2.5-32b on M1 Max via ollama q4_K_M
our signed data: M1 Max · qwen2.5-32b
“is the sky blue?’ runs like this (ollama - m1 max 2e/8p/32 gpu): qwen2.5-14b-instruct-q4_K_M - 26.55 tokens/s qwen2.5-32b-instruct-q4_K_M - 10.62 tokens/s”
- communityconfidence 75%
26.55tok/s — qwen2.5-32b on M1 Max via ollama q4_K_M
our signed data: M1 Max · qwen2.5-32b
“f you tell me the prompt you’re using. ‘why is the sky blue?’ runs like this (ollama - m1 max 2e/8p/32 gpu): qwen2.5-14b-instruct-q4_K_M - 26.55 tokens/s qwen2.5-32b-instruct-q4_K_M - 10.62 tokens/s”
- communityconfidence 60%
11.84tok/s — Qwen2.5-32B on M1 Max via mlx
our signed data: M1 Max · Qwen2.5-32B
“This is great information! You're getting roughly double the speed as my M1 Max 32-core GPU 64GB. I got 11.84 tokens/s on Qwen2.5-32B-Instruct (4bit) MLX”
- communityconfidence 60%
16.00tok/s — Gemma3 27B on M1 Max via mlx
our signed data: M1 Max · Gemma3 27B
“My friend just got MacBook Pro M1 Max 64GB for $1200 used. **Gemma3 27B Q4** on MLX does 16tok/s on that. 800GB/s memory bandwidth. Maybe consider that?”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
- communityconfidence 50%
57.00tok/s — on M1 Max via lm-studio
our signed data: M1 Max
“Max (64GB, 24 GPU)|LM Studio|**17.0** (56.6)|**13.4** (56.8)|**5.9** (54.4)|**38.3** (58.9)| Generation speed is virtually identical (\~54-57 tok/s both). The difference is entirely in prefill: oMLX is up to **10x faster** on long contexts. At 8K context (prefill-test turn 4), L…”
Common questions about M1 Max
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.