Skip to content
llm-speed

Free tool · source benchmarks

LLM speed estimator

Find measured tokens per second for a model and GPU or Mac. Compare separate model sizes, workloads and reported settings before estimating what your setup might deliver. Every displayed rate links to its source.

Compare the reported settings in each result.

gemma3 on RTX 4090

12 measured workload rows from 3 source runs, separated into 12 reported-settings groups. Choose the size, workload and settings closest to yours. Each range describes its included rows; it is not a forecast or confidence interval.

27.4B · chat-short · Q4_K_M46.95tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.

Observed range: 46.95tok/s–46.95tok/s.

27.4B · chat-long · Q4_K_M45.25tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3185 tokens · input/output: 3185/606 tokens.

Observed range: 45.25tok/s–45.25tok/s.

27.4B · concurrent-decode · Q4_K_M46.17tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.

Observed range: 46.17tok/s–46.17tok/s.

27.4B · agent-trace · Q4_K_M45.33tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.

Observed range: 45.33tok/s–45.33tok/s.

12.2B · chat-short · Q4_K_M92.65tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.

Observed range: 92.65tok/s–92.65tok/s.

12.2B · chat-long · Q4_K_M87.60tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3185 tokens · input/output: 3185/881 tokens.

Observed range: 87.60tok/s–87.60tok/s.

12.2B · concurrent-decode · Q4_K_M90.43tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.

Observed range: 90.43tok/s–90.43tok/s.

12.2B · agent-trace · Q4_K_M88.53tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.

Observed range: 88.53tok/s–88.53tok/s.

4.3B · chat-short · Q4_K_M195.0tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 118 tokens · input/output: 118/256 tokens.

Observed range: 195.0tok/s–195.0tok/s.

4.3B · chat-long · Q4_K_M187.6tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3185 tokens · input/output: 3185/585 tokens.

Observed range: 187.6tok/s–187.6tok/s.

4.3B · concurrent-decode · Q4_K_M193.1tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 1024 tokens · input/output: 1024/256 tokens.

Observed range: 193.1tok/s–193.1tok/s.

4.3B · agent-trace · Q4_K_M190.0tok/s observed median · 1 row / 1 run

RTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GB

ollama 0.31.1 · suite suite-v1 · client 0.0.3

Batch: 1 · context: 3934 tokens · input/output: 3934/840 tokens.

Observed range: 190.0tok/s–190.0tok/s.

Hardware benchmarks · Model benchmarks

How to estimate LLM speed from benchmarks

Start with the same model size, quantization, hardware configuration and runtime. Then match the workload, batch size and actual token lengths. A short chat result does not establish speed for a long prompt or several simultaneous users.

Tokens per second and waiting time answer different questions

Decode tok/s measures token generation speed. First-token latency captures the wait before generation starts. Keep both alongside input and output lengths when comparing a chat or coding workflow. Tokenizers differ, so equal token rates do not guarantee equal useful work.

What the groups mean

This lookup includes successful positive decode measurements from matching model and hardware families in up to the latest 500 listed submissions. Groups preserve reported model size, quantization, backend and version, hardware summary, suite/client versions, workload, batch, context and input/output lengths. Missing configuration fields keep separate source runs apart.

Matching reported fields does not prove identical prompts, model files, runtime flags or hardware. The median and minimum/maximum describe the included rows; they do not predict an unmeasured setup. Open the source run and use the methodology to assess comparability.

For a hardware purchase, check the memory estimator and tested model files as well. A fast submitted run alone does not establish that your chosen file and context will fit.

Full benchmark table · Contribute a measurement · Download benchmark data