Skip to content
llm-speed

RTX 4090 vs RTX 5090 for local LLMs: speed and VRAM

The RTX 5090 adds 8 GB of GPU memory over a standard RTX 4090. That can matter when your model and context exceed 24 GB. If your current setup already fits and responds fast enough, an upgrade needs a measured benefit on your workload.

RTX 5090 vs RTX 4090

Individual submitted decode results; model and workload shown below.

NVIDIA Ada/Blackwell
356.1tok/s
decode (individual submitted result)
Model
stable-code-instruct-3b
Reported hardware
RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 3…
Workload / suite
concurrent-decode · suite-v1
Runtime
llama.cpp · version not reported
Quantization field
Not reported
Submitted
2026-07-01
source: r_q9f15lz6831
NVIDIA Ada/Blackwell
195.0tok/s
decode (individual submitted result)
Model
gemma3
Reported hardware
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
Workload / suite
chat-short · suite-v1
Runtime
ollama · 0.31.1
Quantization field
Q4_K_M
Submitted
2026-07-02
source: r_dlanfbgym0h
These independently selected results can use different models, workloads and settings. They do not establish a hardware or model winner.

4090 or 5090: what changes for local LLMs?

NVIDIA reference specifications, checked September 12, 2026.
SpecificationRTX 4090RTX 5090
GPU memory24 GB GDDR6X32 GB GDDR7
ArchitectureAda LovelaceBlackwell
Total graphics power450 W575 W

Sources: NVIDIA RTX 4090 and RTX 5090 specifications. Graphics power is a reference rating, not measured inference consumption or total PC draw. Check the exact board's power, connector and case requirements.

Does 32 GB let you run a model that 24 GB cannot?

Sometimes. Budget for the exact model file, context cache and runtime buffers together. A setup needing more than the available memory on a 4090 might fit fully on a 5090. A setup exceeding the available memory on both still needs a different quantization, shorter context or a measured offload configuration. A parameter count alone cannot settle this.

Our September 5 Qwen3.8-27B Q4_K_M setup on the 5090 reported 19,071 MiB of runtime GPU allocation with 16,384 tokens configured; Gemma 4 12B QAT Q4_0 reported 7,931 MiB. These allocations are not peak whole-device memory or a verified fit test on a 4090. The actual chat inputs reached roughly 3,200 tokens. Inspect the exact model files, settings and repeated measurements, then use the memory-fit estimator to shortlist your own configuration.

How much faster is the 5090 for inference?

The two share-card figures above are independently selected submissions. Different models or workloads can explain their difference; dividing those figures does not establish the benefit of upgrading your GPU. For a useful comparison, match the model artifact, quantization, runtime, prompt length, output length, cache state and concurrent request count.

Measure both first-token wait and generation speed. A document-heavy request can spend much of its time processing the prompt; fast streaming alone does not make the whole request fast. Our 5090 long-context study shows that distinction, but has no matched 4090 run. See the troubleshooting guide before attributing a slow setup to the card.

When does upgrading make sense?

  1. Keep a working 4090 when your chosen model, context and simultaneous requests fit, and task quality and response time meet your needs.
  2. Evaluate a 5090 when the extra memory solves a demonstrated fit problem, or a comparable test shows a speed improvement that matters for the work you actually do.
  3. Compare total upgrade cost using a current local quote, expected resale proceeds and any required PSU or case change. Record actual workload power if electricity cost affects the decision.

Save your current configuration and a representative task before changing hardware. After the change, repeat that task and check successful answers, first-token time and completion time. The measurements on this page cover inference; they do not establish training or fine-tuning speed.

Reviewed September 12, 2026. Reference hardware specifications and submitted benchmark configurations are separate evidence. For recording a comparable run, follow the benchmark setup guide.

Need the long-form table? Open the RTX 5090 vs RTX 4090 comparison for every overlapping (model × hardware) row, source runs, and methodology.