Best hardware for a local coding agent
Pick a rig that runs Qwen3-Coder-Next, Qwen2.5-Coder-32B, gpt-oss, and DeepSeek as a daily-driver coding agent without you waiting on it.
As of today, the fastest submitted coding-agent run is DeepSeek-Coder-V2-Lite-Instruct on RTX 5090 (32GB) at 309.5tok/s. Coverage spans 1 distinct hardware across the coder-tagged catalog — sparse populations should be read as a research lead, not a final answer.
Fastest coder-model run on the leaderboard: DeepSeek-Coder-V2-Lite-Instruct on RTX 5090 (32GB) at 309.5tok/s.
A local coding agent stresses three things at once: enough VRAM (or unified memory) to hold a 14B-32B coder model at a reasonable quantization, prefill speed for long tool-use prompts, and decode speed for the model's reply. Below are the configurations we have benchmark data for, ranked by decode tok/s on the workloads tagged 'coder'. We list every result so you can make the call yourself, not just the headline number.
Submitted benchmarks
| Hardware | Coder model | decode tok/s | Run |
|---|---|---|---|
| RTX 5090 (32GB) | DeepSeek-Coder-V2-Lite-Instruct | 309.5tok/s | r_0_gs1rgl2fl |
Side-by-side comparisons
See also: All hardware · All models · Methodology