Free tool
LLM VRAM calculator
Pick a model, quantization, and rig. For a unified-memory machine, enter its installed memory. We estimate whether it fits in your VRAM or unified memory (weights + KV cache + overhead). Compare the estimate with submitted speed measurements for the model and hardware family where available; the run's quantization, context and runtime may differ.
Pick a model and a rig above. Or browse the cheatsheet of every measured (model × hardware) cell, or the tok/s predictor.
Which GGUF file should I download?
Start with a file your runtime supports, then check memory at your intended context length. These two files were tested on our RTX 5090 on September 5, 2026. They are concrete starting points for comparison, not a quality ranking or a complete list of current models.
Qwen3.8-27B · Q4_K_M
Qwen3.8-27B-Q4_K_M.gguf
18.97 GB file on disk. Our runtime reported 18.62 GiB for model, context and compute buffers in the tested configuration.
View exact Q4_K_M filegemma-4-12b-it-qat · Q4_0
gemma-4-12b-it-qat-q4_0.gguf
6.98 GB file on disk. Our runtime reported 7.75 GiB for model, context and compute buffers in the tested configuration.
View exact Q4_0 fileFile size is not the memory needed to run a model. Our text-only test used llama.cpp CUDA, a 16K configured context and one active request; the actual prompts were shorter. Runtime allocation is not peak whole-device memory. Different context, vision inputs, concurrency, tensor precision or runtime can change requirements. We have not measured these files on a Mac.
Compare measured speed and full settings · Download the benchmark data
For a Mac, enter your installed unified memory in the calculator. Leaving that field blank uses the family's catalog maximum. The estimate reserves about 20% for the OS and other applications; check the exact file, runtime and workload before buying.
How the estimate works
Memory ≈ weights + FP16 KV cache + a fixed runtime allowance. Source-backed model profiles use attention dimensions and exact file sizes where shown; other selections use parameter-count and quantization heuristics. A file size is a planning input, not resident VRAM. All values are decimal GB and 1K means 1,000 tokens. Context includes input and generated tokens for one sequence. For mixture-of-experts models the total parameter count drives memory — all experts must be resident — even though only a few billion are active per token (which is why they decode fast). Apple unified memory reserves ~20% for the OS. These are estimates, not guaranteed upper bounds. Recurrent state, runtime buffers, tensor precision and vision components can change memory use. Allow extra memory for other applications and simultaneous requests. The submitted tok/s describes the linked run, not a speed guarantee for your selected settings. Read the full methodology.