RTX 5090 vs RTX 4090
Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs RTX 4090 — every number links to the run it came from.
Verdict
RTX 4090 has no submitted run yet, so there is no head-to-head number to report. We do have 9 models measured for RTX 5090 (peak 309 tok/s decode on DeepSeek-Coder-V2-Lite-Instruct) — the table below is that baseline. Run the llm-speed suite on RTX 4090 and this page fills in the second column automatically.
RTX 5090 benchmarks — no shared model with RTX 4090 yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| DeepSeek-Coder-V2-Lite-Instruct | 309.5tok/s | chat-short | r_0_gs1rgl2fl |
| Qwen2.5-Coder-7B-Instruct | 244.9tok/s | chat-short | r_mln72x5zbis |
| Qwen3.6-35B-A3B-Q4_K_M.gguf | 224.0tok/s | chat-short | r_g1aw3mizd32 |
| Qwen2.5-Coder-14B-Instruct | 136.0tok/s | chat-short | r_p_f63tcgans |
| Qwen3.6-27B-Q4_K_M.gguf | 73.96tok/s | chat-short | r_f7ulllqu4vk |
| Qwen2.5-Coder-32B-Instruct | 71.91tok/s | chat-short | r_nkbs6d3-d21 |
| Codestral-22B-v0.1 | 69.95tok/s | chat-short | r_sqr8liqh4ii |
| gemma-4-31B-it-Q4_K_M.gguf | 69.67tok/s | chat-short | r_pa7nuerd9tl |
| Qwen2.5-32B-Instruct | 68.71tok/s | chat-short | r_bjy5a5izxjc |
See also: RTX 5090 benchmarks · RTX 4090 benchmarks · All hardware · All models · Methodology