Skip to content
llm-speed

How fast is Gemma 3 on an RTX 4090? (4B, 12B, 27B)

Published 2026-07-03 · updated 2026-09-12

Gemma 3's recorded short-chat decode speeds on a reported 24 GB RTX 4090 were 195.00 tok/s for the 4B variant, 92.65 for 12B and 46.95 for 27B. These are three individual July 2, 2026 submissions using Ollama 0.31.1 and Q4_K_M. They establish performance for the recorded text workloads, not maximum-context fit or a universal quality ranking.

Gemma 3 speed by model size and prompt length

All three source configurations report an AMD EPYC 7443 host with 252 GB system RAM. Short chat used 118 input tokens and 256 output tokens. Long chat used 3,185 input tokens, with different generated output lengths below. Each model has one submission, so these are observations rather than repeat averages.

Decode tokens per second. Links open the source run; long-chat output lengths are shown separately.
VariantShort chatLong chatLong output tokens
4B195.00187.56585
12B92.6587.60881
27B46.9545.25606

Recorded first-token times for 4B, 12B and 27B were 706, 689 and 817 ms on short chat, versus 979, 1,444 and 1,998 ms on long chat. The runtime reports model-loading time as part of each request, so these are not isolated prefill timings or a controlled warm-cache comparison. File revisions, full GPU placement and peak memory are not recorded in these public rows.

Does long context change the streaming rate?

It can. The longer workload here recorded slightly lower decode rates as well as longer first-token waits. Different answer lengths and runtime conditions prevent assigning the whole difference to context. Our earlier claim that long prompts do not change streaming speed was too broad.

Google's Gemma 3 model card lists a 128K context limit for 4B, 12B and 27B. A model capability limit is different from the context your hardware can serve. These roughly 3.2K-input tests do not measure 128K operation or image-input performance.

Ollama's context documentation explains that larger allocations require more memory. Check the actual allocated context and GPU/CPU placement using ollama ps, then test your intended prompt length. Use our slow-model checklist if loading or offload is dominating the wait.

Which size should you try?

Start with a variant that answers your actual tasks well, then compare latency and memory at your intended context. The 4B submission streams faster here; this benchmark does not score answer correctness or prove that 27B is best for every task. A completed run on a GPU-equipped host also does not establish spare VRAM or complete GPU residency.

For a newer measured alternative, our Gemma 4 12B QAT study includes three repetitions and pinned files on an RTX 5090. It uses different hardware and settings, so it is not a matched Gemma 3 versus Gemma 4 result. The memory-fit estimator can help shortlist a configuration before you validate it locally.

Record your own configuration

Use the benchmark setup guide and record the model artifact, runtime, quantization, allocated context and GPU placement. The standard benchmark flow does not reproduce every historical setting automatically.

Reviewed September 12, 2026. Source measurements submitted July 2, 2026; this review is not a new benchmark.