Skip to content
llm-speed

DeepSeek-R1 on an RTX 4090: which size actually fits?

Published 2026-07-02 · updated 2026-09-12

A slow DeepSeek-R1 32B run does not prove that every Q4 configuration needs more than 24 GB. The exact artifact, allocated context, runtime buffers and other resident models determine memory use. Our cited 32B submission is slow, but does not record GPU-layer placement or peak memory. The earlier instruction to buy a larger card was not supported by that evidence.

What our RTX 4090 submissions actually measured

These July 2, 2026 submissions report a 24 GB RTX 4090, AMD EPYC 75F3 host and 504 GB system RAM, using Ollama 0.31.1 and Q4_K_M. The table consistently uses chat-long: 3,142 input tokens and 1,024 generated tokens per row. One run per variant is an observation, not a repeated comparison.

Recorded client decode rate for the long-chat workload; each value links to its source.
Reported model sizeDecode tok/s
8.2B131.95
14.8B81.09
32.8B3.76

The corresponding short-chat rows have no primary decode or first-token measurement. Backend timing fields exist separately; we do not silently substitute them for missing client measurements. Recorded first-token delays on reasoning-model requests can include thinking before visible output, so these results do not isolate prompt processing or time to a correct final answer. The three public fingerprints are empty, and the rows do not pin a model-file revision or prove GPU residency.

Can DeepSeek-R1 32B fit on a 24 GB card?

Check the configuration before deciding. The Ollama 32B listing checked September 12 shows a 32.8B Q4_K_M artifact with a rounded 20 GB download size. Download size is not total runtime allocation: context cache, buffers and other loaded software also need memory. It therefore supplies neither a guaranteed fit nor proof that a 24 GB GPU cannot run the model.

The official DeepSeek release distinguishes distilled models from the full R1 checkpoint. The size labels here refer to the submitted smaller variants; they are not measurements of the full R1 model or of its overall reasoning quality.

Check a slow setup before buying more VRAM

  1. Record the exact model tag or file, quantization, runtime version and allocated context. Start with a context length your task actually needs.
  2. Use ollama ps to inspect the loaded model's CPU/GPU split. Ollama's GPU-placement documentation explains the processor column. Save this evidence while the model is loaded.
  3. Check other resident models and repeat a representative request. Record first visible output, completion time and whether the answer succeeds; a fast token stream can still produce a long or unhelpful answer.

If the intended configuration exceeds available memory, compare a shorter context, another quantization, measured CPU offload or a larger-memory setup. Validate answer quality after changing quantization. The troubleshooting guide walks through the diagnosis; our 4090-to-5090 upgrade checklist helps evaluate the hardware decision once the bottleneck is known.

Contribute a comparable result using the benchmark setup guide. Preserve the source settings alongside the result; signing a submission does not independently verify the machine or reproduce its configuration.

Reviewed September 12, 2026. Measurements submitted July 2, 2026; no new DeepSeek benchmark or maximum-context fit test is claimed.