RTX 5090 vs M3 Ultra
Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs M3 Ultra — every number links to the run it came from.
Verdict
On mlx-community/stable-code-instruct-3b-4bit, the RTX 5090 leads — 356 tok/s decode versus 192 tok/s on the M3 Ultra, 1.8× faster. Across 15 models measured on both, RTX 5090 is faster on 15 of 15. Every cell in the table below links to the submitted run it came from.
512 GB is the historical maximum announced in 2025. Apple's model-specific support page checked September 7, 2026 lists 96 GB configurable to 256 GB. Check the actual machine: a family ceiling is not installed memory, current availability or a measured configuration. 2025 launch specification · Model-specific specification
View M3 Ultra page →Models that have data on both rigs
See also: RTX 5090 benchmarks · M3 Ultra benchmarks · All hardware · All models · Methodology