RTX 4080 SUPER LLM benchmark
No RTX 4080 SUPER LLM benchmarks yet. Be the first to submit a signed, reproducible tok/s run from your own RTX 4080 SUPER.
Will it run local LLMs?
No signed RTX 4080 SUPER run yet, but its 16GB of VRAM tells you what fits. At 4-bit it comfortably runs models up to about 15B parameters with room for context, and up to roughly 23B with tighter quantization. Check a specific model with the VRAM-fit tool.
For measured decode speed on the nearest hardware we have benchmarked: RTX 5090 · RTX 4090 · RTX 3090. Run the suite on your RTX 4080 SUPER to fill this page in.
No RTX 4080 SUPER benchmarks yet.
Run on YOUR hardware to populate this page: pipx install llm-speed && llm-speed bench
$ pipx install llm-speed && llm-speed bench
Community folklore on RTX 4080 SUPER
1 unverified claim extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 60%
24.00tok/s — qwen3-coder on RTX 4080 Super via lm-studio
our signed data: RTX 4080 Super · qwen3-coder
“ing LM Studio. My setup is an NVIDIA RTX 4080 Super (16GB VRAM) with 32GB of system RAM. In the LM Studio chat UI, I was getting about 20-24 tok/s. The speed felt ok and I thought I was on the right track. But I hit a hard wall with the memory bottleneck pretty quickly. The co…”
Common questions about RTX 4080 SUPER
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.