State of the local LLM — July 2026
The canonical answer to “what’s the fastest local LLM right now” as of July 2026, measured under llm-speed suite-v1. Numbers are wall-clock decode tok/s on the highest-decode workload that successfully ran. Every cell links to the run that produced it.
Headline cells
Top 10 decode tok/s — July 2026
Editor’s notes
July 2026 community submissions span NVIDIA GPUs and Apple Silicon. The table ranks the highest successful decode workload for each model and hardware pairing; configurations, quantization and workloads differ, so this is not a controlled hardware comparison. Each headline links to its submitted run. The 70B+ category uses the reported total parameter size, and coder-family speed does not measure coding quality.
182 runs landed in July 2026. To reproduce any number on this page, install the CLI and run the suite on the same model + hardware:
pipx install llm-speed
llm-speed verify
llm-speed benchMethodology: /methodology · Privacy: /privacy · Source: github.com/meadow-kun/llm-speed