Decode vs prefill tok/s: what LLM speed numbers actually mean
Published 2026-07-02 · updated 2026-09-27
Prefill processes the prompt; decode generates the answer. Prefill throughput counts input tokens per second, while decode throughput counts output tokens per second. For a responsive local LLM, check both the wait before the first token and the speed of the stream that follows.
Prefill, decode and time to first token
An autoregressive model can process known prompt tokens in parallel during prefill. During ordinary decode, each new output token depends on earlier ones. At low concurrency, decode is often limited by memory bandwidth; batching, context length, runtime and model architecture can change the bottleneck. There is no universal ratio between the two speeds.
Time to first token (TTFT) measures elapsed time until the first streamed output reaches the measurement point. Prefill contributes to it, but client-observed TTFT can also include queuing, model loading and transport. A warm, idle local server and a busy remote service are different tests. TNG's prefill and decode explanation shows how concurrent requests affect these phases.
Measured: fast streaming after a 20-second wait
On September 6, 2026, we measured Gemma 4 12B IT QAT Q4_0 on an RTX 5090 with llama.cpp CUDA 9725a31. The model was already loaded, with one active request, prompt caching disabled and thinking off. Each input length was measured three times.

September 6 snapshot: a longer first-token wait can accompany fast generation. Separate axes start at zero; three repetitions describe this configuration, not a confidence interval or a quality score.
Download chart (PNG)Download chart (SVG)Download citationAll nine measurements (JSON)
Reuse under CC BY 4.0. Credit llm-speed and link to the study, test date and configuration. Identify any changes you make to the chart.
| Actual input tokens | First token, median (s) | Decode, median (tok/s) |
|---|---|---|
| 23,872 | 3.355 | 136.27 |
| 47,652 | 7.901 | 133.03 |
| 95,212 | 19.901 | 115.98 |
With about four times the input, the median first-token wait grew about sixfold while decode speed fell about 15%. A headline of “116 tok/s” would miss the roughly 20 seconds before streaming began. These are observations for this configuration, not a scaling rule for every model.
Actual outputs were 98, 100 and 96 tokens respectively. Configured context was 131,072 tokens; this was not a full 128K-input test or an answer-quality evaluation. See the full study, observed ranges and settings or download the nine measurements and source-run links.
Which metric should guide your choice?
- Chat and coding: check TTFT, streaming speed and whether the answer is correct. Fast generation alone does not establish a useful coding model.
- Long documents, repositories and RAG: measure the actual input length and cache state you expect. A short-prompt result cannot predict the wait on a large document.
- Batch processing: compare total throughput at the same concurrency, while checking latency for each request. Aggregate tok/s is not one user's streaming speed.
Before buying hardware, compare memory capacity and measured configurations and use the VRAM fit calculator as an estimate. Weight fit alone does not guarantee room for the context or a latency target.
Can you estimate the waiting time?
For an uncached prompt, input tokens divided by a prefill rate measured at a comparable length gives a rough prompt-processing estimate. For example, 10,000 tokens at an assumed 3,300 input tok/s is about 3 seconds. This is arithmetic, not a measured 10K result or a promise of client-observed TTFT. Do not extrapolate a short-prompt rate unchanged to a much longer prompt.
A supported prefix cache can avoid reprocessing a matching prompt prefix; it is different from simply keeping the model loaded. Compare cold, warm and prefix-cached results separately. Shorter relevant inputs may reduce waiting, but removing needed context can harm the answer.
Diagnosing a slow setup? Use the local LLM troubleshooting checklist to check model loading, GPU placement and context before changing hardware.
Does prefix caching remove the first-token wait?
It can move work outside the timed request. In our September 16 exploratory study, Qwen3.8-27B 4-bit ran on a 40-core M5 Max with 48 GiB unified memory, MLX 0.32.1 and mlx-lm 0.31.3. The synthetic document contained 12,263 input tokens. Prepared-prefix trials reused 12,199 tokens and submitted 64 new tokens. Both paths used batch one, thinking off and a 256-token output cap.
| Timing boundary | Fresh context | Prepared prefix |
|---|---|---|
| Timed request to first token | 74.39–133.19 s | 0.71–1.34 s |
| Prefix preparation outside request | None | 118.65–134.77 s |
| Preparation + request to first token | 74.39–133.19 s | 119.51–136.11 s |
The last row adds preparation and first-token time within each trial before taking the range. It excludes model loading and is not total answer time. A sub-second cached request therefore does not mean a new document was processed in under a second. Report preparation separately, and measure repeated reuse in the application before estimating an amortized saving.
This shared-workstation study rebuilt the prefix for each pair; background activity was uncontrolled. Cached and fresh outputs diverged in all three pairs. It does not establish answer equivalence, a general speedup or an app-level cache hit rate, and it is separate from signed suite-v1 results. The full study and pinned model revision and 12 raw measurements preserve the settings and timing boundaries. MLX LM's prompt-cache documentation describes preparing and reusing a cache; those instructions are not a benchmark.
Check the measurement behind a speed claim
Match model artifact and quantization, runtime, actual input and output tokens, concurrency, cache state and warm-up. Report TTFT alongside decode and total request time, with repeated measurements. The benchmark cheatsheet links individual configurations; the methodology explains the suite. A signed result preserves provenance, not a guarantee of identical performance or model quality on another machine.
Reviewed September 27, 2026. The tables preserve the September 6 and September 16 studies; no new benchmark was run for this explainer.