Skip to content
llm-speed

Qwen3-Coder-Next vs Qwen2.5-Coder-32B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen3-Coder-Next vs Qwen2.5-Coder-32B-Instruct — every number links to the run it came from.

Verdict

Qwen3-Coder-Next has no submitted run yet, so there is no head-to-head number to report. We do have 2 hardware configs measured for Qwen2.5-Coder-32B-Instruct (peak 72 tok/s decode on RTX 5090 (32GB)) — the table below is that baseline. Run the llm-speed suite on Qwen3-Coder-Next and this page fills in the second column automatically.

model
Qwen3-Coder-Next
Qwen · 80B-A3B
View Qwen3-Coder-Next page →
model
Qwen2.5-Coder-32B-Instruct
Qwen · 32B
View Qwen2.5-Coder-32B-Instruct page →

Qwen2.5-Coder-32B-Instruct benchmarks — no shared hardware with Qwen3-Coder-Next yet

Hardwaredecode tok/sWorkloadRun
RTX 5090 (32GB)71.91tok/schat-shortr_nkbs6d3-d21
M3 Ultra (60-core GPU)34.48tok/schat-shortr_721b4bls_oq

See also: Qwen3-Coder-Next benchmarks · Qwen2.5-Coder-32B-Instruct benchmarks · All hardware · All models · Methodology