Skip to content
llm-speed

llm-speed

LLM speed benchmarks: community-verified tok/s for every model + hardware combo

The benchmark suite for local + hosted LLM inference.Run it. See your tok/s. Compare your rig.

$ pipx install llm-speed && llm-speed bench

Runs in about a minute. Auto-detects your hardware and backends. Open source.

See a signed 356 tok/s run
80 signed runs3 hosts30 (model × hardware) cellsApache-2.0
Fastest local signed run356tok/sstable-code-instruct-3bon RTX 5090

Latest LLM benchmarks

20 runs · suite-v1

Explore local-LLM speed