Skip to content
llm-speed

RTX 5090 vs H100 SXM

Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 5090 vs H100 SXM — every number links to the run it came from.

Verdict

H100 SXM has no submitted run yet, so there is no head-to-head number to report. We do have 21 models measured for RTX 5090 (peak 356 tok/s decode on stable-code-instruct-3b) — the table below is that baseline. Run the llm-speed suite on H100 SXM and this page fills in the second column automatically.

hardware
RTX 5090
NVIDIA Ada/Blackwell · 32 GB VRAM
View RTX 5090 page →
hardware
H100 SXM
NVIDIA Hopper · 80 GB VRAM
View H100 SXM page →

RTX 5090 benchmarks — no shared model with H100 SXM yet

Modeldecode tok/sWorkloadRun
stable-code-instruct-3b356.1tok/sconcurrent-decoder_q9f15lz6831
gpt-oss-20b318.4tok/schat-shortr_b9ul-vxh9sc
DeepSeek-Coder-V2-Lite-Instruct309.5tok/schat-shortr_0_gs1rgl2fl
Qwen3-Coder-30B-A3B-Instruct259.9tok/schat-shortr_c7qyvvmmsv1
Qwen2.5-Coder-7B-Instruct255.3tok/schat-shortr_2b9o6y_49mi
Qwen2.5-7B-Instruct246.7tok/schat-shortr_3yn-4321hp-
Llama-3.1-8B-Instruct232.2tok/schat-shortr_kfrkg-vn376
Qwen3.6-35B-A3B-Q4_K_M.gguf224.0tok/schat-shortr_g1aw3mizd32
Yi-Coder-9B-Chat199.3tok/schat-shortr_u4iojm6-ekg
gemma-2-9b-it152.7tok/schat-shortr__b_bzmmab_8
gemma-4-12b-it-qat142.6tok/slong-context-decayr_-lh-vczf7ik
phi-4140.9tok/schat-shortr_e-k4aea8ipr
Qwen2.5-Coder-14B-Instruct140.0tok/schat-shortr_opuj21f13-_
Qwen2.5-14B-Instruct133.3tok/schat-shortr_tj9bu7gvnvh
Codestral-22B-v0.1100.3tok/sconcurrent-decoder_4q040m4scic
Qwen3.6-27B-Q4_K_M.gguf73.96tok/schat-shortr_f7ulllqu4vk
Qwen2.5-Coder-32B-Instruct71.91tok/schat-shortr_nkbs6d3-d21
Qwen2.5-32B-Instruct71.90tok/schat-shortr_twfs86tf_xf
gemma-4-31B-it-Q4_K_M.gguf69.67tok/schat-shortr_pa7nuerd9tl
Qwen3-32B69.45tok/schat-shortr_phvxm9dcak0
Qwen3.8-27B65.73tok/schat-shortr_c3ps9wjygi9

See also: RTX 5090 benchmarks · H100 SXM benchmarks · All hardware · All models · Methodology