M1 Ultra LLM benchmark
No M1 Ultra LLM benchmarks yet. Be the first to submit a signed, reproducible tok/s run from your own M1 Ultra.
Will it run local LLMs?
No signed M1 Ultra run yet, but its 128GB of unified memory tells you what fits. At 4-bit it comfortably runs models up to about 115B parameters with room for context, and up to roughly 181B with tighter quantization. Check a specific model with the VRAM-fit tool.
For measured decode speed on the nearest hardware we have benchmarked: M3 Ultra · M4 Max · M3 Pro. Run the suite on your M1 Ultra to fill this page in.
No M1 Ultra benchmarks yet.
Run on YOUR hardware to populate this page: pipx install llm-speed && llm-speed bench
$ pipx install llm-speed && llm-speed bench
Community folklore on M1 Ultra
2 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 60%
12.00tok/s — Qwen2.5-72B on M1 Ultra via lm-studio
our signed data: M1 Ultra · Qwen2.5-72B
“By comparison, my M1 Ultra does about 12 tokens/s for Qwen2.5-72B-Instruct (4bit). The extra bandwidth is just insanely good. BTW One other thing you can try is using speculative decoding”
- communityconfidence 55%
12.70tok/s — on M1 Ultra via mlx fp16
our signed data: M1 Ultra
“nd, vanilla JS frontend, WebSocket streaming **Some benchmarks on M1 Ultra (64GB):** |Model|Speed|Notes| |:-|:-|:-| |GooseOne 2.9B (fp16)|12.7 tok/s|Constant memory, no KV cache| |Z-Image Turbo (Q4)|77s / 1024×1024|Metal acceleration via mflux| The RNN advantage that made me b…”
Common questions about M1 Ultra
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.