Skip to content
llm-speed

RTX 4090 vs DGX Spark

Signed, community-submitted decode tok/s, prefill, and TTFT for RTX 4090 vs DGX Spark — every number links to the run it came from.

Verdict

DGX Spark has no submitted run yet, so there is no head-to-head number to report. We do have 8 models measured for RTX 4090 (peak 195 tok/s decode on gemma3) — the table below is that baseline. Run the llm-speed suite on DGX Spark and this page fills in the second column automatically.

hardware
RTX 4090
NVIDIA Ada/Blackwell · 24 GB VRAM
View RTX 4090 page →
hardware
DGX Spark
NVIDIA GB10 Grace Blackwell · up to 128 GB unified memory
View DGX Spark page →

RTX 4090 benchmarks — no shared model with DGX Spark yet

Modeldecode tok/sWorkloadRun
gemma3195.0tok/schat-shortr_dlanfbgym0h
qwen3-coder179.9tok/sconcurrent-decoder_wjq32z47vlp
qwen2.5-coder161.1tok/sconcurrent-decoder_mv8n8k9wu1e
llama3.1154.4tok/schat-shortr_h1ub_1uxzdh
gpt-oss141.7tok/schat-longr_iu2sfa9ykvw
deepseek-r1133.8tok/sagent-tracer_fg77v2hhohb
glm-4.7-flash129.9tok/sagent-tracer_o636l3cc-rr
qwen3.644.18tok/sagent-tracer_h_659oy695r

See also: RTX 4090 benchmarks · DGX Spark benchmarks · All hardware · All models · Methodology