Skip to content
llm-speed

gemma-4-12b-it-qat on RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

Workload results

WorkloadBackendModeldecode tok/sprefill tok/sTTFTp50p95
long-context-decayllama.cpp@1 (9725a31)gemma-4-12b-it-qatQ4_0142.6tok/s7,115.8tok/s3,355ms7.0ms7.0ms

Long-context results

gemma-4-12b-it-qat · llama.cpp. The workload overview shows the smallest successful point. This table exposes every reported target, including unsuccessful or missing measurements.

Targets estimate tokens from character count. Actual prompt tokens come from the runtime; neither is the configured context capacity. “—” means unmeasured or unavailable.

Scroll the table sideways for all measurements.

Long-context measurements and status for each estimated target.
Estimated targetActual prompt tokensFirst token (s)Decode tok/sOutput tokensStatus
32,00023,8723.355142.6198measured
64,00047,6528.833133.03100measured
128,00095,21220.224123.6596measured

First-token delay includes processing the prompt with the model loaded. These speed measurements do not test answer quality or the maximum usable context.

Part of our Gemma 4 RTX 5090 long-context study: see all three repetitions, observed ranges and the 131,072-context server configuration.

Reproduce on your machine

Start with this model and workload selection. This command does not capture the original artifact, backend version or server settings:

$ pipx install llm-speed && llm-speed bench --model 'gemma-4-12b-it-qat' --workload 'long-context-decay'

Match quantization, context capacity, caching and runtime settings for a useful comparison. Duration varies by workload and hardware; the published client may differ from this run's version. How it's measured.

Embed this run

Drop the badge into a README, blog post, or signature. Each render is a backlink to the signed result.

llm-speed: 143 tok/s on RTX 5090 (32GB) (gemma-4-12b-it-qat)
[![llm-speed: 143 tok/s on RTX 5090 (32GB) (gemma-4-12b-it-qat)](https://llm-speed.com/badge/r_-lh-vczf7ik.svg)](https://llm-speed.com/r/r_-lh-vczf7ik)

Related benchmarks

Provenance

Run ID
r_-lh-vczf7ik
Fingerprint hash
013ca61a09d17996
Public key
0tv44ISLy10gz6Oc6FJig3eWJVLwso9oKemSM4nicKM=
Received
2026-09-06 03:40:00