Skip to content
llm-speed

Best hardware for a local coding agent

Pick a rig that runs Qwen3-Coder-Next, Qwen2.5-Coder-32B, gpt-oss, and DeepSeek as a daily-driver coding agent without you waiting on it.

Verdict

As of today, the fastest submitted coding-agent run is DeepSeek-Coder-V2-Lite-Instruct on RTX 5090 (32GB) at 309.5tok/s. Coverage spans 1 distinct hardware across the coder-tagged catalog — sparse populations should be read as a research lead, not a final answer.

Recommendation

Fastest coder-model run on the leaderboard: DeepSeek-Coder-V2-Lite-Instruct on RTX 5090 (32GB) at 309.5tok/s.

A local coding agent stresses three things at once: enough VRAM (or unified memory) to hold a 14B-32B coder model at a reasonable quantization, prefill speed for long tool-use prompts, and decode speed for the model's reply. Below are the configurations we have benchmark data for, ranked by decode tok/s on the workloads tagged 'coder'. We list every result so you can make the call yourself, not just the headline number.

Submitted benchmarks

HardwareCoder modeldecode tok/sRun
RTX 5090 (32GB)DeepSeek-Coder-V2-Lite-Instruct309.5tok/sr_0_gs1rgl2fl

Side-by-side comparisons

See also: All hardware · All models · Methodology