Skip to content
llm-speed

The best GPU for a local coding agent in 2026

Published 2026-07-02 · updated 2026-09-12

Choose the coding model and workload before choosing the GPU. A useful local agent needs correct edits, supported tool calls, enough memory for the repository context and acceptable waiting time. Decode speed helps describe the experience, but it cannot establish coding quality or the best purchase. The historical submissions below are reference points to test against.

Start with the constraint you need to solve

  • You already own a GPU: run representative coding tasks first. If performance disappoints, use the GPU placement and latency checklist before assuming the hardware needs replacing.
  • Your chosen model fits, but the agent feels slow: compare first-token wait, generation, tool execution and successful task completion. Test with repository-sized inputs and the reasoning settings you intend to use.
  • Your model or context does not fit: estimate weights, context cache and runtime memory with the fit checker, then validate a real configuration. Its capacity assumptions may differ from your machine.
  • You are buying a complete system: compare the actual card or Mac configuration, total system price, power supply, memory and running costs. This guide does not contain verified current prices or a cost-per-task ranking.

What the RTX 3090 and 4090 source runs show

Both July 2, 2026 submissions below used Qwen2.5-Coder 7.6B Q4_K_M and Ollama 0.31.1. The chat-long workload had 3,168 input tokens in each case. These are individual submissions, not a controlled hardware trial.

The second submission's reported 48 GB configuration is not evidence for a standard retail card's memory capacity. Host, output length and other conditions differ. A signature identifies a submission; it does not independently verify the reported hardware. These rows cannot support a universal GPU speedup, a current price comparison or a claim to beat hosted APIs.

Model choice changes both speed and waiting time

Two other July 2 submissions report an RTX 3090 with 24 GB, EPYC 7702P, 252 GB system RAM and Ollama 0.31.1. On chat-short, DeepSeek-Coder-V2 15.7B Q4_0 recorded 189.47 decode tok/s and 0.735 s to first token; Qwen2.5-Coder 14.8B Q4_K_M recorded 69.21 tok/s and 0.326 s. Inputs were 117 and 131 tokens respectively; both generated 256 tokens.

The faster stream started later in these observations. Tokenizers, artifacts and quantizations differ, and the runs do not score coding correctness. Do not infer better answers from model size, architecture or tokens per second.

Where the RTX 5090 and Apple Silicon fit

For newer recorded configurations, inspect the Qwen3.8-27B measurements and Gemma 4 study on our RTX 5090. These are speed and latency observations, not an evaluation of which is today's best coding agent. Check model and runtime support in your coding app before shortlisting either.

For larger memory needs, compare the RTX 5090 and M3 Ultra configurations and the Mac Studio memory checklist. The cited M3 Ultra runs use 96 GB unified memory; they do not establish performance on a 512 GB machine. Desktop Max/Ultra systems are not a portability recommendation.

Make a purchase decision from a useful trial

  1. Select repository tasks with checkable outcomes: an edit that passes tests, a correct tool call or a bug fix you can review.
  2. Record actual model, quantization, context, GPU placement and runtime. Keep first-token time, generation time and tool/test time separate.
  3. Repeat on the candidate setup using the same task and settings where supported. Compare successful outcomes and elapsed time before comparing total cost.

Use the source-run cheatsheet to find configurations and benchmark setup instructions to record a result. The standard benchmark flow does not reproduce every historical artifact or run-specific setting automatically.

Source rows checked September 12, 2026; historical measurements above were submitted July 2. No new benchmark, current price survey or coding-quality ranking was performed for this update.