Qwen2.5-Coder-14B-Instruct vs Qwen3-Coder-30B-A3B-Instruct
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-14B-Instruct vs Qwen3-Coder-30B-A3B-Instruct — every number links to the run it came from.
Verdict
On the RTX 5090 (32GB), Qwen3-Coder-30B-A3B-Instruct decodes at 260 tok/s versus 140 tok/s for Qwen2.5-Coder-14B-Instruct, 1.9× faster. Across 2 hardware configs measured on both, Qwen3-Coder-30B-A3B-Instruct is faster on 2 of 2. Every cell in the table below links to the submitted run it came from.
Hardware with data on both models
| Hardware | Qwen2.5-Coder-14B-Instruct decode | Qwen3-Coder-30B-A3B-Instruct decode | Δ | Source runs |
|---|---|---|---|---|
| RTX 5090 (32GB) | 140.0tok/s | 259.9tok/s | -119.9 | r_opuj21f13-_ · r_c7qyvvmmsv1 |
| M3 Ultra (60-core GPU) | 70.51tok/s | 112.2tok/s | -41.6 | r_l36cijqxq4t · r_fpsca03u2o_ |
See also: Qwen2.5-Coder-14B-Instruct benchmarks · Qwen3-Coder-30B-A3B-Instruct benchmarks · All hardware · All models · Methodology