M3 Ultra vs M5 Max
Signed, community-submitted decode tok/s, prefill, and TTFT for M3 Ultra vs M5 Max — every number links to the run it came from.
Verdict
M5 Max has no submitted run yet, so there is no head-to-head number to report. We do have 1 model measured for M3 Ultra (peak 130 tok/s decode on mlx-community/Llama-3.1-8B-Instruct-4bit) — the table below is that baseline. Run the llm-speed suite on M5 Max and this page fills in the second column automatically.
M3 Ultra benchmarks — no shared model with M5 Max yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| mlx-community/Llama-3.1-8B-Instruct-4bit | 130.2tok/s | chat-short | r_v2pbc0rq2l4 |
See also: M3 Ultra benchmarks · M5 Max benchmarks · All hardware · All models · Methodology