Skip to content
llm-speed

RTX 5090 vs M3 Ultra for local LLMs: speed vs memory

Published 2026-07-02 · updated 2026-09-06

The recorded RTX 5090 configurations generated text faster in the four examples below. The measured M3 Ultra had more memory: 96 GB unified, compared with 32 GB of GPU memory on the 5090. These historical submissions help frame speed and capacity decisions, but they do not isolate the hardware or establish a universal winner.

Recorded short-chat speed and first-token latency

Every cell uses the source run's chat-short workload. The RTX 5090 submissions used llama.cpp on July 1, 2026. The M3 Ultra submissions used a 60-core GPU, 96 GB unified memory and MLX 0.31.3 on April 28, 2026. These are individual measurements, not averages of repeated controlled tests.

Each cell shows decode tokens per second, then time to the first token in milliseconds. Links open the full source run. On a small screen, scroll the table horizontally.

Model familyRTX 5090 · llama.cppM3 Ultra · MLX
Stable Code Instruct 3B341.9 tok/s
174.9 ms first token
192.5 tok/s
226.5 ms first token
DeepSeek Coder V2 Lite Instruct293.1 tok/s
297.0 ms first token
168.3 tok/s
449.5 ms first token
gpt-oss-20b318.4 tok/s
172.5 ms first token
152.7 tok/s
692.1 ms first token
Qwen2.5 Coder 32B Instruct67.4 tok/s
383.5 ms first token
34.5 tok/s
909.4 ms first token

How comparable are these runs?

The workload labels match, but the full configurations do not. The Mac model names identify 4-bit variants (MXFP4-Q4 for gpt-oss); the GPU records omit quantization and backend version. The GPU prompt-token counts are recorded as zero, so they do not establish the actual prompt length. Mac prompt counts range from 127 to 166 tokens. Output lengths are 256 tokens except the GPU gpt-oss run, which produced 253.

Different artifacts, runtimes, tokenizers and submission dates prevent attributing the entire gap to the processor. The results do not measure answer quality, noise, electricity use or total ownership cost. A signed submission preserves its reported result; it is not independent proof of the hardware configuration.

Memory capacity: what was tested and what was offered

The Mac in this table had 96 GB, not 512 GB. Apple's March 2025 launch announcement described M3 Ultra configurations with up to 512 GB. The Apple UK technical specification checked September 6, 2026 lists 96 GB configurable to 256 GB. Historical capacity, current listed configurations and the machine actually measured are different facts; confirm the exact configuration available before buying.

Model weights, context cache and runtime buffers all consume memory. More capacity can let you load a larger model or context, but these short-chat runs do not show its speed or quality. Use the memory-fit estimator as a starting estimate, then test the intended configuration. The Mac Studio model guide separates its historical measurements from capacity assumptions.

How to use this evidence for your setup

First choose a model that does your task well, then check its memory needs at your intended context length. For a configuration that fits, compare first-token delay, generation speed and task success using the same model artifact and workload wherever possible. If it does not fit on one GPU, measure the proposed offload or larger-memory setup rather than extrapolating from this table.

For newer downloadable models, see our September 5 RTX 5090 studies of Qwen3.8-27B and Gemma 4 12B IT QAT. Those include three repetitions, pinned artifacts, tested settings and memory allocation. We have not measured those configurations on an M3 Ultra, so they do not fill the missing cross-platform comparison.

Add a comparable measurement

Follow the benchmark methodology and record the model artifact, quantization, runtime version, context length and cache settings. The standard install command starts the benchmark flow; it does not reproduce all settings in these historical submissions automatically.

$ pipx install llm-speed && llm-speed bench

Corrected September 6, 2026: the earlier version mixed different workloads and overstated a hardware-wide speed advantage. This table uses chat-short throughout. The source measurements remain dated April 28 and July 1; this is an article update, not a new hardware benchmark.