LLM hardware buying guide
Pick a guide that matches what you actually want to do. Every recommendation is anchored to a real submitted run on the llm-speed leaderboard — no affiliate fluff, just numbers you can click through.
Best hardware for a local coding agent
Pick a rig that runs Qwen3-Coder-Next, Qwen2.5-Coder-32B, gpt-oss, and DeepSeek as a daily-driver coding agent without you waiting on it.
Best rig for Qwen3-Coder-Next
Qwen3-Coder-Next is an 80B-parameter MoE with ~3B active. The activation pattern means it punches above its weight on Apple Silicon and sane consumer GPUs. Here's the data we have.
Fastest rig for Qwen2.5-Coder-32B (local)
Qwen2.5-Coder-32B is the strongest fully-open dense coding model that fits a single prosumer card at 4-bit. Here's the fastest decode tok/s submitted across every GPU and Apple Silicon tier.
Fastest rig for Qwen3-Coder-30B-A3B (local)
Qwen3-Coder-30B-A3B is a 30B mixture-of-experts with ~3B active — it decodes like a small model but codes like a big one. Here's the fastest decode tok/s submitted across every GPU and Apple Silicon tier.
Local vs hosted: when does buying a GPU pay off?
At low usage, hosted APIs win on $/Mtok. At high sustained usage, a 4090 or M3 Ultra wins. Here's the break-even math, run against live numbers.
Best Mac for running local LLMs
M-series chips trade off bandwidth, GPU cores, and unified memory ceiling. Here's the data, ranked by decode tok/s on Qwen-class models.
Best GPU for local LLMs under $2,000
RTX 4090, RTX 5080, used 3090, RX 7900 XTX, Arc B580. Here's where each lands on real workloads.
Best GPU for local LLMs in 2026
RTX 5090, 4090, used 3090, 5070 Ti and the rest — ranked on real decode tok/s across the models people actually run locally. No affiliate picks, just submitted runs.
Cheapest rig that runs a 70B model comfortably
Llama-3.3-70B and Qwen2.5-72B at 4-bit need ~40 GB of memory. Here's the minimum-spec hardware that holds the model and still serves usable decode tok/s.
Fastest GPU for Llama 3.3 70B (local, 2026)
Llama 3.3 70B has become the default open-weights 70B baseline. Here's the fastest decode tok/s submitted across every GPU and Apple Silicon class on the leaderboard.
Apple Silicon vs NVIDIA for local LLMs
Unified memory vs VRAM, MLX vs CUDA, M-series Ultra vs RTX 5090. Here's how the two architectures actually compare on real submitted runs.
Fastest small-MoE coder on a Mac (Qwen3-Coder-30B, DeepSeek-Coder-V2-Lite)
Two MoE coders — Qwen3-Coder-30B-A3B and DeepSeek-Coder-V2-Lite — both clear 100 tok/s decode on an M3 Ultra at 4-bit. Here's how they rank across every Apple Silicon tier we have data for.