Qwen3-Coder-Next vs Qwen2.5-Coder-32B-Instruct
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen3-Coder-Next vs Qwen2.5-Coder-32B-Instruct — every number links to the run it came from.
Verdict
Qwen3-Coder-Next has no submitted run yet, so there is no head-to-head number to report. We do have 2 hardware configs measured for Qwen2.5-Coder-32B-Instruct (peak 72 tok/s decode on RTX 5090 (32GB)) — the table below is that baseline. Run the llm-speed suite on Qwen3-Coder-Next and this page fills in the second column automatically.
Qwen2.5-Coder-32B-Instruct benchmarks — no shared hardware with Qwen3-Coder-Next yet
| Hardware | decode tok/s | Workload | Run |
|---|---|---|---|
| RTX 5090 (32GB) | 71.91tok/s | chat-short | r_nkbs6d3-d21 |
| M3 Ultra (60-core GPU) | 34.48tok/s | chat-short | r_721b4bls_oq |
See also: Qwen3-Coder-Next benchmarks · Qwen2.5-Coder-32B-Instruct benchmarks · All hardware · All models · Methodology