Local vs hosted: when does buying a GPU pay off?
At low usage, hosted APIs win on $/Mtok. At high sustained usage, a 4090 or M3 Ultra wins. Here's the break-even math, run against live numbers.
The reference local rig on the leaderboard is RTX 5090 (32GB) at 318.4tok/s. At ~$0.40 per million output tokens (typical 70B hosted price), a $2k GPU breaks even on output cost alone after roughly 5 billion output tokens generated — power and duty cycle move that number in either direction. The per-rig table below lets you re-do the arithmetic for your own usage.
Reference local rig on the leaderboard: RTX 5090 (32GB) at 318.4tok/s.
We don't sell hardware and we don't take affiliate commissions on hosted APIs, so the framing is just arithmetic. A consumer GPU's break-even point against a hosted endpoint depends on three things: your sustained decode tok/s, the hosted price per million output tokens, and your duty cycle. Below is a comparison table with each row anchored to a real submitted benchmark.
Submitted benchmarks
| Hardware | Model | decode tok/s | Run |
|---|---|---|---|
| RTX 5090 (32GB) | gpt-oss-20b | 318.4tok/s | r_b9ul-vxh9sc |
| RTX 4090 (24GB) | gemma3 | 195.0tok/s | r_dlanfbgym0h |
| RTX 3090 (24GB) | deepseek-coder-v2 | 189.5tok/s | r_o2-1w665rtq |
| M3 Ultra (60-core GPU) | mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit | 168.3tok/s | r_l_v1-zq_qaz |
| RTX 4090 (48GB) | qwen2.5-coder | 161.1tok/s | r_mv8n8k9wu1e |
| M4 Max (40-core GPU) | qwen3-coder-bench-32k | 113.3tok/s | r_roktphpc--8 |
| M3 Pro (18-core GPU) | mlx-community-Qwen2.5-7B-Instruct-4bit | 30.52tok/s | r_llzv_g-ymaf |
Side-by-side comparisons
See also: All hardware · All models · Methodology