Qwen2.5-Coder-32B-Instruct vs Codestral-22B-v0.1
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-32B-Instruct vs Codestral-22B-v0.1 — every number links to the run it came from.
How to read these results
Selected submissions for RTX 5090 (32GB) report 72 tok/s on Qwen2.5-Coder-32B-Instruct and 100 tok/s on Codestral-22B-v0.1. The table selects each side's fastest submitted row independently, grouped by hardware label. Matching labels do not establish matching model artifacts, quantization, backend, workload or context. These peaks are reference results, not a controlled speed or answer-quality ranking. Inspect both source runs before choosing hardware.
Hardware with data on both models
| Hardware | Qwen2.5-Coder-32B-Instruct decode | Codestral-22B-v0.1 decode | Δ | Source runs |
|---|---|---|---|---|
| RTX 5090 (32GB) | 71.91tok/s chat-short · llama.cpp · quantization unreported | 100.3tok/s concurrent-decode · llama.cpp · quantization unreported | -28.4 | r_nkbs6d3-d21 · r_4q040m4scic |
| M3 Ultra (60-core GPU) | 34.48tok/s chat-short · mlx · quantization unreported | 47.49tok/s chat-short · mlx · quantization unreported | -13.0 | r_721b4bls_oq · r_79dvtag5fd_ |
See also: Qwen2.5-Coder-32B-Instruct benchmarks · Codestral-22B-v0.1 benchmarks · All hardware · All models · Methodology