State of the local LLM — July 2026
The canonical answer to “what’s the fastest local LLM right now” as of July 2026, measured under llm-speed suite-v1. Numbers are wall-clock decode tok/s on the highest-decode workload that successfully ran. Every cell links to the run that produced it.
Headline cells
Top 10 decode tok/s — July 2026
Editor’s notes
July 2026. The leaderboard now holds over 200 signed runs, the large majority measured this month across the RTX 5090, RTX 4090, and Apple Silicon. The fastest decode is a small model on a top GPU (stable-code-3b on an RTX 5090), while among capable larger models the small mixture-of-experts coders lead: DeepSeek-Coder-V2-Lite and Qwen3-Coder-30B-A3B both clear 260 tok/s on the 5090, far ahead of dense models of similar total size. Every headline below links to its signed run under suite-v1.
49 runs landed in July 2026. To reproduce any number on this page, install the CLI and run the suite on the same model + hardware:
pipx install llm-speed
llm-speed verify
llm-speed benchMethodology: /methodology · Privacy: /privacy · Source: github.com/meadow-kun/llm-speed