Skip to content
llm-speed

RTX 5090 vs M3 Ultra

Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs M3 Ultra — every number links to the run it came from.

Verdict

On mlx-community/stable-code-instruct-3b-4bit, the RTX 5090 leads — 356 tok/s decode versus 192 tok/s on the M3 Ultra, 1.8× faster. Across 15 models measured on both, RTX 5090 is faster on 15 of 15. Every cell in the table below links to the submitted run it came from.

hardware
RTX 5090
NVIDIA Ada/Blackwell · 32 GB VRAM
View RTX 5090 page →
hardware
M3 Ultra
Apple M3 · up to 512 GB unified memory

512 GB is the historical maximum announced in 2025. Apple's model-specific support page checked September 7, 2026 lists 96 GB configurable to 256 GB. Check the actual machine: a family ceiling is not installed memory, current availability or a measured configuration. 2025 launch specification · Model-specific specification

View M3 Ultra page →

Models that have data on both rigs

See also: RTX 5090 benchmarks · M3 Ultra benchmarks · All hardware · All models · Methodology