Qwen2.5-Coder-32B-Instruct vs Codestral-22B-v0.1
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-32B-Instruct vs Codestral-22B-v0.1 — every number links to the run it came from.
Verdict
On the RTX 5090 (32GB), Codestral-22B-v0.1 decodes at 100 tok/s versus 72 tok/s for Qwen2.5-Coder-32B-Instruct, 1.4× faster. Across 2 hardware configs measured on both, Codestral-22B-v0.1 is faster on 2 of 2. Every cell in the table below links to the submitted run it came from.
Hardware with data on both models
| Hardware | Qwen2.5-Coder-32B-Instruct decode | Codestral-22B-v0.1 decode | Δ | Source runs |
|---|---|---|---|---|
| RTX 5090 (32GB) | 71.91tok/s | 100.3tok/s | -28.4 | r_nkbs6d3-d21 · r_4q040m4scic |
| M3 Ultra (60-core GPU) | 34.48tok/s | 47.49tok/s | -13.0 | r_721b4bls_oq · r_79dvtag5fd_ |
See also: Qwen2.5-Coder-32B-Instruct benchmarks · Codestral-22B-v0.1 benchmarks · All hardware · All models · Methodology