Skip to content
llm-speed

RTX 4090D LLM benchmark

No RTX 4090D LLM benchmarks yet. Be the first to submit a signed, reproducible tok/s run from your own RTX 4090D.

Will it run local LLMs?

No signed RTX 4090D run yet, but its 24GB of VRAM tells you what fits. At 4-bit it comfortably runs models up to about 23B parameters with room for context, and up to roughly 37B with tighter quantization. Check a specific model with the VRAM-fit tool.

For measured decode speed on the nearest hardware we have benchmarked: RTX 5090 · RTX 4090 · RTX 3090. Run the suite on your RTX 4090D to fill this page in.

No RTX 4090D benchmarks yet.

Run on YOUR hardware to populate this page: pipx install llm-speed && llm-speed bench

$ pipx install llm-speed && llm-speed bench

Common questions about RTX 4090D

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the RTX 4090D FAQ →