gemma-4-12b-it-qat on RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB
Workload results
| Workload | Backend | Model | decode tok/s | prefill tok/s | TTFT | p50 | p95 |
|---|---|---|---|---|---|---|---|
| long-context-decay | llama.cpp@1 (9725a31) | gemma-4-12b-it-qatQ4_0 | 134.2tok/s | 6,022.4tok/s | 3,964ms | 7.4ms | 7.4ms |
Long-context results
gemma-4-12b-it-qat · llama.cpp. The workload overview shows the smallest successful point. This table exposes every reported target, including unsuccessful or missing measurements.
Targets estimate tokens from character count. Actual prompt tokens come from the runtime; neither is the configured context capacity. “—” means unmeasured or unavailable.
Scroll the table sideways for all measurements.
| Estimated target | Actual prompt tokens | First token (s) | Decode tok/s | Output tokens | Status |
|---|---|---|---|---|---|
| 32,000 | 23,872 | 3.964 | 134.18 | 98 | measured |
| 64,000 | 47,652 | 7.901 | 131.71 | 100 | measured |
| 128,000 | 95,212 | 19.765 | 115.05 | 96 | measured |
First-token delay includes processing the prompt with the model loaded. These speed measurements do not test answer quality or the maximum usable context.
Part of our Gemma 4 RTX 5090 long-context study: see all three repetitions, observed ranges and the 131,072-context server configuration.
Reproduce on your machine
Start with this model and workload selection. This command does not capture the original artifact, backend version or server settings:
$ pipx install llm-speed && llm-speed bench --model 'gemma-4-12b-it-qat' --workload 'long-context-decay'
Match quantization, context capacity, caching and runtime settings for a useful comparison. Duration varies by workload and hardware; the published client may differ from this run's version. How it's measured.
Embed this run
Drop the badge into a README, blog post, or signature. Each render is a backlink to the signed result.
[](https://llm-speed.com/r/r_8q-uf-rq-gp)Related benchmarks
- More gemma-4-12b-it-qat benchmarks — every backend and rig that has run this model.
- More RTX 5090 (32GB) LLM benchmarks — every model measured on this hardware.
Provenance
- Run ID
- r_8q-uf-rq-gp
- Fingerprint hash
- 013ca61a09d17996
- Public key
- 0tv44ISLy10gz6Oc6FJig3eWJVLwso9oKemSM4nicKM=
- Received
- 2026-09-06 03:38:47