M3 Pro vs M3 Max
Signed, community-submitted decode tok/s, prefill, and TTFT for M3 Pro vs M3 Max — every number links to the run it came from.
Verdict
M3 Max has no submitted run yet, so there is no head-to-head number to report. We do have 9 models measured for M3 Pro (peak 287 tok/s decode on mlx-community-Qwen2.5-0.5B-Instruct-4bit) — the table below is that baseline. Run the llm-speed suite on M3 Max and this page fills in the second column automatically.
M3 Pro benchmarks — no shared model with M3 Max yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| mlx-community-Qwen2.5-0.5B-Instruct-4bit | 286.5tok/s | chat-short | r_akcbpx5vcqa |
| mlx-community-Qwen2.5-7B-Instruct-4bit | 30.52tok/s | chat-short | r_llzv_g-ymaf |
| mlx-community/Llama-3.1-8B-Instruct-4bit | 29.20tok/s | chat-short | r_h0-use1ypnb |
| mlx-community/gemma-2-9b-it-4bit | 21.97tok/s | chat-short | r_q7t7b7dcuz5 |
| mlx-community/stable-code-instruct-3b-4bit | 19.37tok/s | chat-short | r_pqjsvd-cub4 |
| mlx-community/phi-4-4bit | 15.61tok/s | chat-short | r_-w8hnn61va_ |
| mlx-community-Qwen2.5-14B-Instruct-4bit | 14.41tok/s | chat-short | r_ia73dzeue0b |
| mlx-community/Qwen3-32B-4bit | 7.16tok/s | chat-short | r_pnrrpcdqfo4 |
| mlx-community/Qwen2.5-32B-Instruct-4bit | 7.05tok/s | chat-short | r_f4x3xaan2ay |
See also: M3 Pro benchmarks · M3 Max benchmarks · All hardware · All models · Methodology