Skip to content
llm-speed

Free tool · source benchmarks

LLM speed estimator

Find measured tokens per second for a model and GPU or Mac. Compare separate model sizes, workloads and reported settings before estimating what your setup might deliver. Every displayed rate links to its source.

Compare the reported settings in each result.

Qwen3.8-27B on RTX 5090

6 measured workload rows from 3 source runs, separated into 6 reported-settings groups. Choose the size, workload and settings closest to yours. Each range describes its included rows; it is not a forecast or confidence interval.

Size not reported · chat-short · Q4_K_M65.66tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 114 tokens · input/output: 114/256 tokens.

Observed range: 65.66tok/s–65.66tok/s.

Size not reported · chat-long · Q4_K_M65.10tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 3184 tokens · input/output: 3184/874 tokens.

Observed range: 65.10tok/s–65.10tok/s.

Size not reported · chat-short · Q4_K_M65.73tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 114 tokens · input/output: 114/256 tokens.

Observed range: 65.73tok/s–65.73tok/s.

Size not reported · chat-long · Q4_K_M65.09tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 3184 tokens · input/output: 3184/874 tokens.

Observed range: 65.09tok/s–65.09tok/s.

Size not reported · chat-short · Q4_K_M65.43tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 114 tokens · input/output: 114/256 tokens.

Observed range: 65.43tok/s–65.43tok/s.

Size not reported · chat-long · Q4_K_M65.10tok/s observed median · 1 row / 1 run

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

llama.cpp 1 (9725a31) · suite suite-v1 · client 0.0.5

Batch: 1 · context: 3184 tokens · input/output: 3184/874 tokens.

Observed range: 65.10tok/s–65.10tok/s.

Hardware benchmarks · Model benchmarks

How to estimate LLM speed from benchmarks

Start with the same model size, quantization, hardware configuration and runtime. Then match the workload, batch size and actual token lengths. A short chat result does not establish speed for a long prompt or several simultaneous users.

Tokens per second and waiting time answer different questions

Decode tok/s measures token generation speed. First-token latency captures the wait before generation starts. Keep both alongside input and output lengths when comparing a chat or coding workflow. Tokenizers differ, so equal token rates do not guarantee equal useful work.

What the groups mean

This lookup includes successful positive decode measurements from matching model and hardware families in up to the latest 500 listed submissions. Groups preserve reported model size, quantization, backend and version, hardware summary, suite/client versions, workload, batch, context and input/output lengths. Missing configuration fields keep separate source runs apart.

Matching reported fields does not prove identical prompts, model files, runtime flags or hardware. The median and minimum/maximum describe the included rows; they do not predict an unmeasured setup. Open the source run and use the methodology to assess comparability.

For a hardware purchase, check the memory estimator and tested model files as well. A fast submitted run alone does not establish that your chosen file and context will fit.

Full benchmark table · Contribute a measurement · Download benchmark data