Skip to content
llm-speed

RTX 5090 (32GB) LLM benchmark

The fastest LLM measured on the RTX 5090 (32GB) is stable-code-instruct-3b at 356.1 decode tok/s via llama.cpp (signed run). Across 103 reproducible runs on 21 models, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.

Fastest known config on RTX 5090 (32GB)

356.1 decode tok/s

stable-code-instruct-3b via llama.cpp. see full run

Original measurements · 5 September 2026

Gemma 4 12B vs Qwen3.8-27B on RTX 5090

Compare response speed, first-token delay and GPU allocation on the same rig. Both tested model files fit entirely on this 32 GB GPU with a 16,384-token context configured. The prompts below are shorter than that configured limit.

Gemma 4 12B IT QAT

Q4_0 · 6.98 GB model file

7.75 GiB runtime GPU allocation

Model + context + compute buffers; all 49 layers on GPU.

Settings and three source runs →

Qwen3.8-27B

Q4_K_M · 18.97 GB model file

18.62 GiB runtime GPU allocation

Model + context + compute buffers; all 65 layers on GPU.

Settings and three source runs →
Medians of three repetitions per model and workload. Parentheses show observed minimum–maximum, not confidence intervals. Scroll the table on small screens.
Model / workloadDecode tok/s ↑First token ms ↓Actual input tokensActual output tokens
Gemma 4 12B IT QATQ4_0 · chat-short139.5(137.3141.1)59.7(52.261.8)121256
Qwen3.8-27BQ4_K_M · chat-short65.7(65.465.7)117.5(117.1131.7)114256
Gemma 4 12B IT QATQ4_0 · chat-long135.6(135.6137.1)619.3(577.2622.9)3,188620
Qwen3.8-27BQ4_K_M · chat-long65.1(65.165.1)909.4(908.0910.5)3,184874

Same llama.cpp CUDA runtime, suite-v1 text prompts, one request at a time, warm models, thinking and prompt caching off. Output caps: 256 tokens for chat-short and 1,024 for chat-long. Tokenizers and actual output lengths differ.

Use these results to shortlist models for your latency and memory needs, then test answer quality on your own tasks. Coding accuracy, reasoning quality, vision and multi-user capacity were not measured. Runtime allocation is not peak whole-device memory; other idle services remained GPU-resident.

RTX 5090 model studies

Repeated tests with pinned model files, runtime settings, observed ranges and source runs. Choose a study for configuration details beyond the leaderboard's single result.

gemma-4-12b-it-qat

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
long-context-decayllama.cpp@1 (9725a31)Q4_0142.6tok/s7,115.8tok/s3,355msr_-lh-vczf7ik
long-context-decayllama.cpp@1 (9725a31)Q4_0136.3tok/s7,161.0tok/s3,334msr_1srqmpjnkgd
long-context-decayllama.cpp@1 (9725a31)Q4_0134.2tok/s6,022.4tok/s3,964msr_8q-uf-rq-gp
chat-shortllama.cpp@1 (9725a31)Q4_0141.1tok/s2,027.9tok/s59.7msr_v5arqmazf31
chat-longllama.cpp@1 (9725a31)Q4_0135.6tok/s5,147.7tok/s619msr_v5arqmazf31
chat-shortllama.cpp@1 (9725a31)Q4_0139.5tok/s1,958.9tok/s61.8msr_c2podhe9qwc
chat-longllama.cpp@1 (9725a31)Q4_0137.1tok/s5,118.0tok/s623msr_c2podhe9qwc
chat-shortllama.cpp@1 (9725a31)Q4_0137.3tok/s2,316.4tok/s52.2msr_8xmm65n1fhq
chat-longllama.cpp@1 (9725a31)Q4_0135.6tok/s5,523.3tok/s577msr_8xmm65n1fhq

Qwen3.8-27B

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp@1 (9725a31)Q4_K_M65.66tok/s969.8tok/s118msr_dz5a3lzxhrj
chat-longllama.cpp@1 (9725a31)Q4_K_M65.10tok/s3,496.8tok/s911msr_dz5a3lzxhrj
chat-shortllama.cpp@1 (9725a31)Q4_K_M65.73tok/s973.6tok/s117msr_c3ps9wjygi9
chat-longllama.cpp@1 (9725a31)Q4_K_M65.09tok/s3,506.5tok/s908msr_c3ps9wjygi9
chat-shortllama.cpp@1 (9725a31)Q4_K_M65.43tok/s865.3tok/s132msr_7mod38qsldj
chat-longllama.cpp@1 (9725a31)Q4_K_M65.10tok/s3,501.1tok/s909msr_7mod38qsldj

Coder-V2-Lite-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-293.1tok/sno data297msr_bfpto9so2o1
chat-longllama.cpp-181.2tok/sno data361msr_bfpto9so2o1
concurrent-decodellama.cpp-268.4tok/sno datano datar_bfpto9so2o1
agent-tracellama.cpp-204.5tok/s24,978.7tok/s80.7msr_bfpto9so2o1
chat-shortllama.cpp-309.5tok/sno data268msr_0_gs1rgl2fl

gpt-oss-20b

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-318.4tok/sno data173msr_b9ul-vxh9sc
chat-longllama.cpp-301.6tok/sno data239msr_b9ul-vxh9sc
concurrent-decodellama.cpp-302.8tok/sno datano datar_b9ul-vxh9sc
agent-tracellama.cpp-304.5tok/s26,907.5tok/s77.7msr_b9ul-vxh9sc
chat-shortllama.cpp-69.42tok/sno data333msr_r9h57uts9lr

Yi-Coder-9B-Chat

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-66.79tok/sno data317msr_zlh6az5q0o_
chat-longllama.cpp-69.60tok/sno data1,195msr_zlh6az5q0o_
concurrent-decodellama.cpp-70.99tok/sno datano datar_zlh6az5q0o_
agent-tracellama.cpp-71.76tok/s4,478.0tok/s397msr_zlh6az5q0o_
chat-shortllama.cpp-199.3tok/sno data36.2msr_u4iojm6-ekg

stable-code-instruct-3b

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-341.9tok/sno data175msr_q9f15lz6831
chat-longllama.cpp-289.1tok/sno data232msr_q9f15lz6831
concurrent-decodellama.cpp-356.1tok/sno datano datar_q9f15lz6831
agent-tracellama.cpp-292.8tok/s44,387.5tok/s42.7msr_q9f15lz6831
chat-shortllama.cpp-331.0tok/sno data36.7msr_8l57cim1i10

Qwen2.5-32B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-68.71tok/sno data77.0msr_bjy5a5izxjc
chat-longllama.cpp-65.36tok/sno data968msr_bjy5a5izxjc
concurrent-decodellama.cpp-68.60tok/sno datano datar_bjy5a5izxjc
agent-tracellama.cpp-66.45tok/s10,407.9tok/s214msr_bjy5a5izxjc
chat-shortllama.cpp-71.90tok/sno data115msr_twfs86tf_xf

Qwen2.5-14B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-130.3tok/sno data48.6msr_xr4qdv1hgf2
chat-longllama.cpp-122.7tok/sno data489msr_xr4qdv1hgf2
concurrent-decodellama.cpp-131.4tok/sno datano datar_xr4qdv1hgf2
agent-tracellama.cpp-122.8tok/s20,315.4tok/s111msr_xr4qdv1hgf2
chat-shortllama.cpp-133.3tok/sno data55.9msr_tj9bu7gvnvh

Llama-3.1-8B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-229.4tok/sno data35.6msr_qr4srge34da
chat-longllama.cpp-211.4tok/sno data266msr_qr4srge34da
concurrent-decodellama.cpp-223.0tok/sno datano datar_qr4srge34da
agent-tracellama.cpp-210.3tok/s36,720.3tok/s60.8msr_qr4srge34da
chat-shortllama.cpp-232.2tok/sno data33.8msr_kfrkg-vn376

Qwen2.5-7B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-232.7tok/sno data37.0msr_1shiviswt3d
chat-longllama.cpp-232.4tok/sno data239msr_1shiviswt3d
concurrent-decodellama.cpp-240.7tok/sno datano datar_1shiviswt3d
agent-tracellama.cpp-232.5tok/s38,007.7tok/s57.1msr_1shiviswt3d
chat-shortllama.cpp-246.7tok/sno data33.9msr_3yn-4321hp-

Qwen3-32B

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-66.51tok/sno data158msr_-txe_hiq44n
chat-longllama.cpp-62.30tok/sno data1,067msr_-txe_hiq44n
concurrent-decodellama.cpp-66.64tok/sno datano datar_-txe_hiq44n
agent-tracellama.cpp-64.34tok/s8,394.1tok/s259msr_-txe_hiq44n
chat-shortllama.cpp-69.45tok/sno data327msr_phvxm9dcak0

phi-4

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-139.1tok/sno data51.1msr_k1u_k1j_1i2
chat-longllama.cpp-133.7tok/sno data433msr_k1u_k1j_1i2
concurrent-decodellama.cpp-140.2tok/sno datano datar_k1u_k1j_1i2
agent-tracellama.cpp-124.8tok/s21,162.4tok/s102msr_k1u_k1j_1i2
chat-shortllama.cpp-140.9tok/sno data56.7msr_e-k4aea8ipr

gemma-2-9b-it

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-69.45tok/sno data325msr_1_xl4zb5-xj
chat-longllama.cpp-69.42tok/sno data1,044msr_1_xl4zb5-xj
concurrent-decodellama.cpp-68.57tok/sno datano datar_1_xl4zb5-xj
agent-tracellama.cpp-68.46tok/s4,816.8tok/s420msr_1_xl4zb5-xj
chat-shortllama.cpp-152.7tok/sno data58.9msr__b_bzmmab_8

Qwen2.5-Coder-32B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-67.41tok/sno data384msr_v983y0y3r2u
chat-longllama.cpp-69.15tok/sno data1,251msr_v983y0y3r2u
concurrent-decodellama.cpp-70.96tok/sno datano datar_v983y0y3r2u
agent-tracellama.cpp-68.56tok/s4,677.1tok/s414msr_v983y0y3r2u
chat-shortllama.cpp-71.91tok/sno data71.0msr_nkbs6d3-d21

Qwen2.5-Coder-14B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-140.0tok/sno data42.1msr_opuj21f13-_
chat-longllama.cpp-123.0tok/sno data457msr_opuj21f13-_
concurrent-decodellama.cpp-133.8tok/sno datano datar_opuj21f13-_
agent-tracellama.cpp-130.5tok/s21,985.9tok/s104msr_opuj21f13-_
chat-shortllama.cpp-136.0tok/sno data51.3msr_p_f63tcgans

Qwen2.5-Coder-7B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-255.3tok/sno data35.4msr_2b9o6y_49mi
chat-longllama.cpp-244.0tok/sno data228msr_2b9o6y_49mi
concurrent-decodellama.cpp-254.0tok/sno datano datar_2b9o6y_49mi
agent-tracellama.cpp-229.9tok/s38,735.2tok/s57.4msr_2b9o6y_49mi
chat-shortllama.cpp-244.9tok/sno data36.1msr_mln72x5zbis

Qwen3-Coder-30B-A3B-Instruct

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-68.41tok/sno data334msr_pm_a1uf2ufc
chat-longllama.cpp-70.55tok/sno data1,204msr_pm_a1uf2ufc
concurrent-decodellama.cpp-68.04tok/sno datano datar_pm_a1uf2ufc
agent-tracellama.cpp-66.09tok/s4,816.7tok/s451msr_pm_a1uf2ufc
chat-shortllama.cpp-259.9tok/sno data218msr_c7qyvvmmsv1

Codestral-22B-v0.1

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-93.72tok/sno data68.6msr_4q040m4scic
chat-longllama.cpp-87.36tok/sno data849msr_4q040m4scic
concurrent-decodellama.cpp-100.3tok/sno datano datar_4q040m4scic
agent-tracellama.cpp-93.44tok/s10,351.0tok/s191msr_4q040m4scic
chat-shortllama.cpp-69.95tok/sno data199msr_sqr8liqh4ii

Qwen3.6-35B-A3B-Q4_K_M.gguf

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-223.3tok/sno data263msr_dp6sr2_iwcf
chat-shortllama.cpp-224.0tok/sno data167msr_a56-wxl21lk
chat-shortllama.cpp-216.5tok/sno data185msr_k070lz99uzi

gemma-4-31B-it-Q4_K_M.gguf

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-67.37tok/sno data180msr_uut6m_v6ui9
chat-shortllama.cpp-67.27tok/sno data344msr_23fireoga9y

Qwen3.6-27B-Q4_K_M.gguf

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-72.38tok/sno data167msr_yluotk909p8
chat-shortllama.cpp-72.48tok/sno data172msr_v2vrkr4uah1
chat-shortllama.cpp-69.56tok/sno data227msr_u4wa_y_y3vt

Community folklore on RTX 5090 (32GB)

134 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.

  • communityconfidence 75%

    10.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    95fa52dbf8ebec6acaf0105e1e9 Hey all, Just a quick one in case it saves someone else a headache. I was getting really poor throughput (\~10 tok/sec) with Qwen3-Coder-Next-Q4\_K\_S.gguf on llama.cpp, like “this can’t be right” levels, and eventually found a set of args that fix…

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    26.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    ~26 tok/sec with Unsloth Qwen3-Coder-Next-Q4_K_S on RTX 5090 (Windows/llama.cpp) https://preview.redd.it/9gfytpz5srhg1.png?width=692&format=png&auto=w

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    10.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    95fa52dbf8ebec6acaf0105e1e9 Hey all, Just a quick one in case it saves someone else a headache. I was getting really poor throughput (\~10 tok/sec) with Qwen3-Coder-Next-Q4\_K\_S.gguf on llama.cpp, like “this can’t be right” levels, and eventually found a set of args that fix…

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    26.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    ~26 tok/sec with Unsloth Qwen3-Coder-Next-Q4_K_S on RTX 5090 (Windows/llama.cpp) https://preview.redd.it/9gfytpz5srhg1.png?width=692&format=png&auto=w

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    10.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    95fa52dbf8ebec6acaf0105e1e9 Hey all, Just a quick one in case it saves someone else a headache. I was getting really poor throughput (\~10 tok/sec) with Qwen3-Coder-Next-Q4\_K\_S.gguf on llama.cpp, like “this can’t be right” levels, and eventually found a set of args that fix…

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    26.00tok/s Qwen3-Coder-Next on RTX 5090 via llama.cpp Q4_K_S

    our signed data: RTX 5090 · Qwen3-Coder-Next

    ~26 tok/sec with Unsloth Qwen3-Coder-Next-Q4_K_S on RTX 5090 (Windows/llama.cpp) https://preview.redd.it/9gfytpz5srhg1.png?width=692&format=png&auto=w

    source: Reddit · u/Spiritual_Tie_5574 · 2026-02-06

  • communityconfidence 75%

    207.9tok/s Qwen3-Coder on RTX 5090 via sglang AWQ

    our signed data: RTX 5090 · Qwen3-Coder

    hoosing the Framework **RTX 5090 — Qwen3-Coder-30B-A3B-Instruct-AWQ** |Metric|vLLM|SGLang| |:-|:-|:-| |Output throughput|**555.82 tok/s**|207.93 tok/s| |Mean TTFT|**549 ms**|1,558 ms| |Median TPOT|**7.06 ms**|18.84 ms| vLLM wins by 2.7x. SGLang is required `--quantization moe_…

    source: Reddit · u/NoVibeCoding · 2026-03-06

  • communityconfidence 75%

    555.8tok/s Qwen3-Coder on RTX 5090 via sglang AWQ

    our signed data: RTX 5090 · Qwen3-Coder

    atency? # 1. Choosing the Framework **RTX 5090 — Qwen3-Coder-30B-A3B-Instruct-AWQ** |Metric|vLLM|SGLang| |:-|:-|:-| |Output throughput|**555.82 tok/s**|207.93 tok/s| |Mean TTFT|**549 ms**|1,558 ms| |Median TPOT|**7.06 ms**|18.84 ms| vLLM wins by 2.7x. SGLang is required `--qu…

    source: Reddit · u/NoVibeCoding · 2026-03-06

  • communityconfidence 75%

    207.9tok/s Qwen3-Coder on RTX 5090 via sglang AWQ

    our signed data: RTX 5090 · Qwen3-Coder

    hoosing the Framework **RTX 5090 — Qwen3-Coder-30B-A3B-Instruct-AWQ** |Metric|vLLM|SGLang| |:-|:-|:-| |Output throughput|**555.82 tok/s**|207.93 tok/s| |Mean TTFT|**549 ms**|1,558 ms| |Median TPOT|**7.06 ms**|18.84 ms| vLLM wins by 2.7x. SGLang is required `--quantization moe_…

    source: Reddit · u/NoVibeCoding · 2026-03-06

  • communityconfidence 75%

    555.8tok/s Qwen3-Coder on RTX 5090 via sglang AWQ

    our signed data: RTX 5090 · Qwen3-Coder

    atency? # 1. Choosing the Framework **RTX 5090 — Qwen3-Coder-30B-A3B-Instruct-AWQ** |Metric|vLLM|SGLang| |:-|:-|:-| |Output throughput|**555.82 tok/s**|207.93 tok/s| |Mean TTFT|**549 ms**|1,558 ms| |Median TPOT|**7.06 ms**|18.84 ms| vLLM wins by 2.7x. SGLang is required `--qu…

    source: Reddit · u/NoVibeCoding · 2026-03-06

See all 134 claims for RTX 5090 (32GB)

Models measured on RTX 5090 (32GB)

Common questions about RTX 5090 (32GB)

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the RTX 5090 (32GB) FAQ →