Qwen3.8-27B vs Gemma 4 12B: RTX 5090 benchmarks
Published 2026-09-12
Gemma 4 12B QAT generated tokens faster and used less runtime GPU memory in our tested RTX 5090 configurations. This comparison covers measured speed, first-token latency and memory use for the two downloadable models. Coding quality needs a separate evaluation on your tasks.
Short-chat median decode was 139.50 tok/s for Gemma Q4_0 and 65.66 tok/s for Qwen Q4_K_M. Median first-token wait was 59.7 ms and 117.5 ms respectively. The table below is a September 12 snapshot of six accepted September 5 runs, with three repetitions of each model and both workloads.
Measured speed and latency on the same RTX 5090
| Model / workload | Decode tok/s | First token ms | Input tokens | Output tokens |
|---|---|---|---|---|
| Gemma 4 12B QAT Q4_0chat-short | 139.50(137.27–141.12) | 59.7(52.2–61.8) | 121 | 256 |
| Qwen3.8-27B Q4_K_Mchat-short | 65.66(65.43–65.73) | 117.5(117.1–131.7) | 114 | 256 |
| Gemma 4 12B QAT Q4_0chat-long | 135.64(135.61–137.11) | 619.3(577.2–622.9) | 3,188 | 620 |
| Qwen3.8-27B Q4_K_Mchat-long | 65.10(65.09–65.10) | 909.4(908.0–910.5) | 3,184 | 874 |
Both used the same physical rig and llama.cpp CUDA revision 9725a31, suite-v1 text prompts, one active request, warm models, thinking off and prompt caching disabled. Configured context was 16,384 tokens; output caps were 256 for short chat and 1,024 for long chat. Temperature was zero and seed 42. Other idle GPU services remained resident.
The tokenizers and generated answers differ. Long chat produced 620 Gemma tokens versus 874 Qwen tokens, so neither output tok/s nor a shorter completion proves that an equally useful task finished sooner. These roughly 3.2K-token inputs also do not establish performance at the configured 16K limit.
Inspect all three repetitions and settings for Qwen3.8-27B and Gemma 4 12B QAT, or download the 12 measured workload rows as CSV and data, ranges and configuration as JSON.
Which files were tested, and how much memory did they use?
- Qwen3.8-27B Q4_K_M: 18.97 GB model file; runtime logs reported 19,071 MiB (18.62 GiB) for model, context and compute buffers, with all 65 layers on GPU. Exact tested Qwen artifact.
- Gemma 4 12B IT QAT Q4_0: 6.98 GB model file; runtime logs reported 7,931 MiB (7.75 GiB), with all 49 layers on GPU. Exact tested Gemma artifact.
File sizes use decimal GB; allocation uses binary GiB. Runtime allocation is not peak whole-device memory or a guarantee of fit on a smaller GPU. Both configurations fitted this 32 GB card, with the settings above. Changing context, quantization, runtime or concurrency changes the question. Use the memory-fit estimator for an initial shortlist and validate the actual setup.
Which should you try for coding or an agent?
If both answer your tasks adequately, the tested Gemma configuration is a useful starting point when response speed and memory headroom matter. If you prefer Qwen's answers or tool use on your workload, these speed results quantify part of that trade-off; they do not invalidate that choice.
Evaluate repository edits, tool calls and task completion with the exact artifact and settings you will use. Qwen's official model card describes thinking enabled by default and its own evaluation settings; our run explicitly disabled thinking. Google's Gemma model card documents separate capability evaluations. Publisher scores are not a matched quality evaluation of these two quantized, thinking-off setups.
Does the faster model stay fast on long documents?
That requires a separate test. Our Gemma long-context study measured actual inputs up to 95,212 tokens with a different context allocation. We do not have a matching Qwen curve here. Read prefill versus decode to see why fast streaming can follow a long wait.
Use or cite this comparison
Cite the September 5 measurement date, exact quantizations, runtime and tested token lengths alongside this page. The downloadable snapshot includes every contributing run URL and configuration notes. A signed submission preserves provenance; it does not independently certify quality.
The isolated benchmark client included token-usage and quantization parsing fixes. The model pages document the server and client limitations; installing the public CLI alone is not a guarantee of exact end-to-end reproduction. For a first contribution, use the benchmark setup guide.
Published September 12, 2026. Measurements collected September 5, 2026.