Skip to content
llm-speed

RTX 5090 vs M4 Max

Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs M4 Max — every number links to the run it came from.

How to read these results

RTX 5090 and M4 Max both have submitted runs, but no single model has been measured on both yet — so there is no apples-to-apples row. Each side's measured decode tok/s is listed separately below (peaks of 356 and 113 tok/s). Submit a shared workload to turn this into a direct comparison.

hardware
RTX 5090
NVIDIA Ada/Blackwell · 32 GB VRAM
View RTX 5090 page →
hardware
M4 Max
Apple M4 · up to 128 GB unified memory
View M4 Max page →

RTX 5090 benchmarks — no shared model with M4 Max yet

Modeldecode tok/sWorkloadRun
stable-code-instruct-3b356.1tok/sconcurrent-decoder_q9f15lz6831
gpt-oss-20b318.4tok/schat-shortr_b9ul-vxh9sc
DeepSeek-Coder-V2-Lite-Instruct309.5tok/schat-shortr_0_gs1rgl2fl
Qwen3-Coder-30B-A3B-Instruct259.9tok/schat-shortr_c7qyvvmmsv1
Qwen2.5-Coder-7B-Instruct255.3tok/schat-shortr_2b9o6y_49mi
Qwen2.5-7B-Instruct246.7tok/schat-shortr_3yn-4321hp-
Llama-3.1-8B-Instruct232.2tok/schat-shortr_kfrkg-vn376
Qwen3.6-35B-A3B-Q4_K_M.gguf224.0tok/schat-shortr_g1aw3mizd32
Yi-Coder-9B-Chat199.3tok/schat-shortr_u4iojm6-ekg
gemma-2-9b-it152.7tok/schat-shortr__b_bzmmab_8
gemma-4-12b-it-qat142.6tok/slong-context-decayr_-lh-vczf7ik
phi-4140.9tok/schat-shortr_e-k4aea8ipr
Qwen2.5-Coder-14B-Instruct140.0tok/schat-shortr_opuj21f13-_
Qwen2.5-14B-Instruct133.3tok/schat-shortr_tj9bu7gvnvh
Codestral-22B-v0.1100.3tok/sconcurrent-decoder_4q040m4scic
Qwen3.6-27B-Q4_K_M.gguf73.96tok/schat-shortr_f7ulllqu4vk
Qwen2.5-Coder-32B-Instruct71.91tok/schat-shortr_nkbs6d3-d21
Qwen2.5-32B-Instruct71.90tok/schat-shortr_twfs86tf_xf
gemma-4-31B-it-Q4_K_M.gguf69.67tok/schat-shortr_pa7nuerd9tl
Qwen3-32B69.45tok/schat-shortr_phvxm9dcak0
Qwen3.8-27B65.73tok/schat-shortr_c3ps9wjygi9

M4 Max benchmarks — no shared model with RTX 5090 yet

Modeldecode tok/sWorkloadRun
qwen3-coder-bench-32k113.3tok/schat-shortr_roktphpc--8
qwen3-coder109.9tok/sconcurrent-decoder_r0di2hkku1h

See also: RTX 5090 benchmarks · M4 Max benchmarks · All hardware · All models · Methodology