Skip to content
llm-speed

M3 Max vs M3 Ultra

Signed, community-submitted decode tok/s, prefill, and TTFT for M3 Max vs M3 Ultra — every number links to the run it came from.

Verdict

M3 Max has no submitted run yet, so there is no head-to-head number to report. We do have 21 models measured for M3 Ultra (peak 192 tok/s decode on mlx-community/stable-code-instruct-3b-4bit) — the table below is that baseline. Run the llm-speed suite on M3 Max and this page fills in the second column automatically.

hardware
M3 Max
Apple M3 · up to 128 GB unified memory
View M3 Max page →
hardware
M3 Ultra
Apple M3 · up to 192 GB unified memory
View M3 Ultra page →

M3 Ultra benchmarks — no shared model with M3 Max yet

Modeldecode tok/sWorkloadRun
mlx-community/stable-code-instruct-3b-4bit192.5tok/schat-shortr_y2_5y8oo97d
mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit168.3tok/schat-shortr_l_v1-zq_qaz
mlx-community/gpt-oss-20b-MXFP4-Q4152.7tok/schat-shortr_3ijun8ltjnb
mlx-community-Qwen2.5-7B-Instruct-4bit139.6tok/schat-shortr_5r6rhiynenc
mlx-community/Qwen2.5-Coder-7B-Instruct-4bit138.6tok/schat-shortr_uoehjq0nvc0
mlx-community/Llama-3.1-8B-Instruct-4bit130.2tok/schat-shortr_v2pbc0rq2l4
mlx-community/granite-8b-code-instruct-4bit112.3tok/schat-shortr_bue3bee0gw7
mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit112.2tok/schat-shortr_fpsca03u2o_
mlx-community/Yi-Coder-9B-Chat-4bit103.5tok/schat-shortr_3hvui9a1yuc
mlx-community/gemma-2-9b-it-4bit89.45tok/schat-shortr_iz137eqvuzy
lmstudio-community-Qwen3-Next-80B-A3B-Instruct-MLX-4bit80.34tok/schat-shortr_1pl79r50ofy
mlx-community/phi-4-4bit74.37tok/schat-shortr_sqzp-0rdez-
mlx-community-Qwen2.5-14B-Instruct-4bit70.85tok/schat-shortr_v4bq1sviz4o
mlx-community/Qwen2.5-Coder-14B-Instruct-4bit70.51tok/schat-shortr_l36cijqxq4t
mlx-community/starcoder2-15b-4bit64.16tok/schat-shortr_wsxml_39dh_
mlx-community/Codestral-22B-v0.1-4bit47.49tok/schat-shortr_79dvtag5fd_
mlx-community-Qwen2.5-32B-Instruct-4bit34.60tok/schat-shortr_njgxtgyym1e
mlx-community/Qwen2.5-Coder-32B-Instruct-4bit34.48tok/schat-shortr_721b4bls_oq
mlx-community/Qwen3-32B-4bit34.41tok/schat-shortr_anmmc80-aoq
mlx-community/llama-3.3-70b-Instruct-4bit16.78tok/schat-shortr_sx3a4y9n-m4
mlx-community/qwen2.5-72b-Instruct-4bit16.31tok/schat-shortr_5c80gthqlh6

See also: M3 Max benchmarks · M3 Ultra benchmarks · All hardware · All models · Methodology