Skip to content
llm-speed

State of the local LLM

State of the local LLM — July 2026

The canonical answer to “what’s the fastest local LLM right now” as of July 2026, measured under llm-speed suite-v1. Numbers are wall-clock decode tok/s on the highest-decode workload that successfully ran. Every cell links to the run that produced it.

Headline cells

Fastest local model
429.2tok/s
Qwen3-0.6B-4bit on M3 Ultra (60-core GPU) + 96GB unified (mlx)
View run →
Fastest local 70B+ class
—
no 70B-class submissions this month
Fastest local coding agent
309.5tok/s
Coder-V2-Lite-Instruct on RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB (llama.cpp)
View run →

Top 10 decode tok/s — July 2026

Editor’s notes

July 2026 community submissions span NVIDIA GPUs and Apple Silicon. The table ranks the highest successful decode workload for each model and hardware pairing; configurations, quantization and workloads differ, so this is not a controlled hardware comparison. Each headline links to its submitted run. The 70B+ category uses the reported total parameter size, and coder-family speed does not measure coding quality.

182 runs landed in July 2026. To reproduce any number on this page, install the CLI and run the suite on the same model + hardware:

pipx install llm-speed
llm-speed verify
llm-speed bench

Methodology: /methodology · Privacy: /privacy · Source: github.com/meadow-kun/llm-speed