llm-speed
LLM speed benchmarks: community-verified tok/s for every model + hardware combo
The benchmark suite for local + hosted LLM inference.
Run it. See your tok/s. Compare your rig.
$ pipx install llm-speed && llm-speed benchRuns in about a minute. Auto-detects your hardware and backends. Open source.
See a signed 356 tok/s runLatest LLM benchmarks
20 runs · suite-v1- RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
- RTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GB
- RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
- RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
- RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
- RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
- RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
- RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
- RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
- RTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GB
- RTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GB
- RTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GB
- RTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GB
- RTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GB
- RTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GB
- RTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GB
- RTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GB
- RTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GB
- RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB
- RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB
Explore local-LLM speed
Will it run — and how fast?Check any model + quant against your GPU or Mac, with the real measured tok/s.Buying guidesBest GPU, Mac, or rig for local LLMs — ranked on signed runs, no affiliate picks.Head-to-headRTX 5090 vs 4090, Apple vs NVIDIA, model vs model — decode tok/s side by side.Speed cheatsheetDecode tok/s for every model × hardware cell we have data for.Community reports vs signed2,176 tok/s claims from r/LocalLLaMA & HN, next to our measured numbers.FAQHow fast is usable? Does quantization cost speed? Why is long context slow?BlogLong-form on local-LLM speed, grounded in signed benchmark runs.