Skip to content
llm-speed

RTX 5090 vs M4 Max

Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs M4 Max — every number links to the run it came from.

Verdict

M4 Max has no submitted run yet, so there is no head-to-head number to report. We do have 19 models measured for RTX 5090 (peak 356 tok/s decode on stable-code-instruct-3b) — the table below is that baseline. Run the llm-speed suite on M4 Max and this page fills in the second column automatically.

hardware
RTX 5090
NVIDIA Ada/Blackwell · 32 GB VRAM
View RTX 5090 page →
hardware
M4 Max
Apple M4 · up to 128 GB unified memory
View M4 Max page →

RTX 5090 benchmarks — no shared model with M4 Max yet

Modeldecode tok/sWorkloadRun
stable-code-instruct-3b356.1tok/sconcurrent-decoder_q9f15lz6831
gpt-oss-20b318.4tok/schat-shortr_b9ul-vxh9sc
DeepSeek-Coder-V2-Lite-Instruct293.1tok/schat-shortr_bfpto9so2o1
Qwen3-Coder-30B-A3B-Instruct259.9tok/schat-shortr_c7qyvvmmsv1
Qwen2.5-Coder-7B-Instruct255.3tok/schat-shortr_2b9o6y_49mi
Qwen2.5-7B-Instruct240.7tok/sconcurrent-decoder_1shiviswt3d
Llama-3.1-8B-Instruct229.4tok/schat-shortr_qr4srge34da
Qwen3.6-35B-A3B-Q4_K_M.gguf224.0tok/schat-shortr_g1aw3mizd32
phi-4140.2tok/sconcurrent-decoder_k1u_k1j_1i2
Qwen2.5-Coder-14B-Instruct140.0tok/schat-shortr_opuj21f13-_
Qwen2.5-14B-Instruct131.4tok/sconcurrent-decoder_xr4qdv1hgf2
Codestral-22B-v0.1100.3tok/sconcurrent-decoder_4q040m4scic
Qwen3.6-27B-Q4_K_M.gguf73.56tok/schat-shortr_az687c4xasr
Yi-Coder-9B-Chat71.76tok/sagent-tracer_zlh6az5q0o_
Qwen2.5-Coder-32B-Instruct70.96tok/sconcurrent-decoder_v983y0y3r2u
gemma-2-9b-it69.45tok/schat-shortr_1_xl4zb5-xj
gemma-4-31B-it-Q4_K_M.gguf68.77tok/schat-shortr_epzfs8k8ohh
Qwen2.5-32B-Instruct68.71tok/schat-shortr_bjy5a5izxjc
Qwen3-32B66.64tok/sconcurrent-decoder_-txe_hiq44n

See also: RTX 5090 benchmarks · M4 Max benchmarks · All hardware · All models · Methodology