Skip to content
llm-speed

Qwen2.5-Coder-32B-Instruct vs Qwen3-Coder-30B-A3B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-32B-Instruct vs Qwen3-Coder-30B-A3B-Instruct — every number links to the run it came from.

Verdict

On the RTX 5090 (32GB), Qwen3-Coder-30B-A3B-Instruct decodes at 260 tok/s versus 72 tok/s for Qwen2.5-Coder-32B-Instruct, 3.6× faster. Across 2 hardware configs measured on both, Qwen3-Coder-30B-A3B-Instruct is faster on 2 of 2. Every cell in the table below links to the submitted run it came from.

model
Qwen2.5-Coder-32B-Instruct
Qwen · 32B
View Qwen2.5-Coder-32B-Instruct page →
model
Qwen3-Coder-30B-A3B-Instruct
Qwen · 30B-A3B
View Qwen3-Coder-30B-A3B-Instruct page →

Hardware with data on both models

HardwareQwen2.5-Coder-32B-Instruct decodeQwen3-Coder-30B-A3B-Instruct decodeΔSource runs
RTX 5090 (32GB)71.91tok/s259.9tok/s-188.0r_nkbs6d3-d21 · r_c7qyvvmmsv1
M3 Ultra (60-core GPU)34.48tok/s112.2tok/s-77.7r_721b4bls_oq · r_fpsca03u2o_

See also: Qwen2.5-Coder-32B-Instruct benchmarks · Qwen3-Coder-30B-A3B-Instruct benchmarks · All hardware · All models · Methodology