Free tool · source benchmarks
LLM speed estimator
Find measured tokens per second for a model and GPU or Mac. Compare separate model sizes, workloads and reported settings before estimating what your setup might deliver. Every displayed rate links to its source.
Choose a model and hardware above, or start with a measured example below.
How to estimate LLM speed from benchmarks
Start with the same model size, quantization, hardware configuration and runtime. Then match the workload, batch size and actual token lengths. A short chat result does not establish speed for a long prompt or several simultaneous users.
Try a measured configuration
Tokens per second and waiting time answer different questions
Decode tok/s measures token generation speed. First-token latency captures the wait before generation starts. Keep both alongside input and output lengths when comparing a chat or coding workflow. Tokenizers differ, so equal token rates do not guarantee equal useful work.
What the groups mean
This lookup includes successful positive decode measurements from matching model and hardware families in up to the latest 500 listed submissions. Groups preserve reported model size, quantization, backend and version, hardware summary, suite/client versions, workload, batch, context and input/output lengths. Missing configuration fields keep separate source runs apart.
Matching reported fields does not prove identical prompts, model files, runtime flags or hardware. The median and minimum/maximum describe the included rows; they do not predict an unmeasured setup. Open the source run and use the methodology to assess comparability.
For a hardware purchase, check the memory estimator and tested model files as well. A fast submitted run alone does not establish that your chosen file and context will fit.
Full benchmark table · Contribute a measurement · Download benchmark data