Free tool · source benchmarks
LLM speed estimator
Find measured tokens per second for a model and GPU or Mac. Compare separate model sizes, workloads and reported settings before estimating what your setup might deliver. Every displayed rate links to its source.
gemma3 on RTX 4090
12 measured workload rows from 3 source runs, separated into 12 reported-settings groups. Choose the size, workload and settings closest to yours. Each range describes its included rows; it is not a forecast or confidence interval.
27.4B · chat-short · Q4_K_M46.95tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.
Observed range: 46.95tok/s–46.95tok/s.
- r_x23y_sg24pm46.95tok/s · first token 817ms
27.4B · chat-long · Q4_K_M45.25tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3185 tokens · input/output: 3185/606 tokens.
Observed range: 45.25tok/s–45.25tok/s.
- r_x23y_sg24pm45.25tok/s · first token 1,998ms
27.4B · concurrent-decode · Q4_K_M46.17tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.
Observed range: 46.17tok/s–46.17tok/s.
- r_x23y_sg24pm46.17tok/s · first token no data
27.4B · agent-trace · Q4_K_M45.33tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.
Observed range: 45.33tok/s–45.33tok/s.
- r_x23y_sg24pm45.33tok/s · first token 956ms
12.2B · chat-short · Q4_K_M92.65tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.
Observed range: 92.65tok/s–92.65tok/s.
- r_3kmrc135e0e92.65tok/s · first token 689ms
12.2B · chat-long · Q4_K_M87.60tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3185 tokens · input/output: 3185/881 tokens.
Observed range: 87.60tok/s–87.60tok/s.
- r_3kmrc135e0e87.60tok/s · first token 1,444ms
12.2B · concurrent-decode · Q4_K_M90.43tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.
Observed range: 90.43tok/s–90.43tok/s.
- r_3kmrc135e0e90.43tok/s · first token no data
12.2B · agent-trace · Q4_K_M88.53tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.
Observed range: 88.53tok/s–88.53tok/s.
- r_3kmrc135e0e88.53tok/s · first token 815ms
4.3B · chat-short · Q4_K_M195.0tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.
Observed range: 195.0tok/s–195.0tok/s.
- r_dlanfbgym0h195.0tok/s · first token 706ms
4.3B · chat-long · Q4_K_M187.6tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3185 tokens · input/output: 3185/585 tokens.
Observed range: 187.6tok/s–187.6tok/s.
- r_dlanfbgym0h187.6tok/s · first token 979ms
4.3B · concurrent-decode · Q4_K_M193.1tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.
Observed range: 193.1tok/s–193.1tok/s.
- r_dlanfbgym0h193.1tok/s · first token no data
4.3B · agent-trace · Q4_K_M190.0tok/s observed median · 1 row / 1 run
RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB
ollama 0.31.1 · suite suite-v1 · client 0.0.3
Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.
Observed range: 190.0tok/s–190.0tok/s.
- r_dlanfbgym0h190.0tok/s · first token 715ms
How to estimate LLM speed from benchmarks
Start with the same model size, quantization, hardware configuration and runtime. Then match the workload, batch size and actual token lengths. A short chat result does not establish speed for a long prompt or several simultaneous users.
Try a measured configuration
Tokens per second and waiting time answer different questions
Decode tok/s measures token generation speed. First-token latency captures the wait before generation starts. Keep both alongside input and output lengths when comparing a chat or coding workflow. Tokenizers differ, so equal token rates do not guarantee equal useful work.
What the groups mean
This lookup includes successful positive decode measurements from matching model and hardware families in up to the latest 500 listed submissions. Groups preserve reported model size, quantization, backend and version, hardware summary, suite/client versions, workload, batch, context and input/output lengths. Missing configuration fields keep separate source runs apart.
Matching reported fields does not prove identical prompts, model files, runtime flags or hardware. The median and minimum/maximum describe the included rows; they do not predict an unmeasured setup. Open the source run and use the methodology to assess comparability.
For a hardware purchase, check the memory estimator and tested model files as well. A fast submitted run alone does not establish that your chosen file and context will fit.
Full benchmark table · Contribute a measurement · Download benchmark data