RTX 5090 vs RTX 3090
Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs RTX 3090 — every number links to the run it came from.
Verdict
RTX 5090 and RTX 3090 both have submitted runs, but no single model has been measured on both yet — so there is no apples-to-apples row. Each side's measured decode tok/s is listed separately below (peaks of 247 and 189 tok/s). Submit a shared workload to turn this into a direct comparison.
RTX 5090 benchmarks — no shared model with RTX 3090 yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| Qwen2.5-7B-Instruct | 246.7tok/s | chat-short | r_3yn-4321hp- |
| Qwen2.5-Coder-7B-Instruct | 244.9tok/s | chat-short | r_mln72x5zbis |
| Llama-3.1-8B-Instruct | 232.2tok/s | chat-short | r_kfrkg-vn376 |
| Qwen3.6-35B-A3B-Q4_K_M.gguf | 224.0tok/s | chat-short | r_a56-wxl21lk |
| Qwen2.5-Coder-14B-Instruct | 136.0tok/s | chat-short | r_p_f63tcgans |
| Qwen3.6-27B-Q4_K_M.gguf | 73.31tok/s | chat-short | r_wax_x2ryqhk |
| Qwen2.5-Coder-32B-Instruct | 70.96tok/s | concurrent-decode | r_v983y0y3r2u |
| Codestral-22B-v0.1 | 69.95tok/s | chat-short | r_sqr8liqh4ii |
| Qwen3-32B | 69.45tok/s | chat-short | r_phvxm9dcak0 |
| gemma-2-9b-it | 69.45tok/s | chat-short | r_1_xl4zb5-xj |
| gpt-oss-20b | 69.42tok/s | chat-short | r_r9h57uts9lr |
| gemma-4-31B-it-Q4_K_M.gguf | 67.37tok/s | chat-short | r_moawmhigq5d |
RTX 3090 benchmarks — no shared model with RTX 5090 yet
| Model | decode tok/s | Workload | Run |
|---|---|---|---|
| deepseek-coder-v2 | 189.5tok/s | chat-short | r_o2-1w665rtq |
| qwen2.5-coder | 69.21tok/s | chat-short | r_6ahy-dq0f_0 |
See also: RTX 5090 benchmarks · RTX 3090 benchmarks · All hardware · All models · Methodology