Qwen2.5-Coder-7B-Instruct vs Qwen2.5-7B-Instruct
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-7B-Instruct vs Qwen2.5-7B-Instruct — every number links to the run it came from.
Verdict
On the RTX 5090 (32GB), Qwen2.5-Coder-7B-Instruct decodes at 255 tok/s versus 247 tok/s for Qwen2.5-7B-Instruct. Across 2 hardware configs measured on both, Qwen2.5-Coder-7B-Instruct is faster on 1 of 2. Every cell in the table below links to the submitted run it came from.
Hardware with data on both models
| Hardware | Qwen2.5-Coder-7B-Instruct decode | Qwen2.5-7B-Instruct decode | Δ | Source runs |
|---|---|---|---|---|
| RTX 5090 (32GB) | 255.3tok/s | 246.7tok/s | +8.6 | r_2b9o6y_49mi · r_3yn-4321hp- |
| M3 Ultra (60-core GPU) | 138.6tok/s | 139.6tok/s | -1.0 | r_uoehjq0nvc0 · r_5r6rhiynenc |
See also: Qwen2.5-Coder-7B-Instruct benchmarks · Qwen2.5-7B-Instruct benchmarks · All hardware · All models · Methodology