Skip to content
llm-speed

Qwen2.5-Coder-7B-Instruct vs Qwen2.5-7B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-7B-Instruct vs Qwen2.5-7B-Instruct — every number links to the run it came from.

Verdict

On the RTX 5090 (32GB), Qwen2.5-Coder-7B-Instruct decodes at 255 tok/s versus 247 tok/s for Qwen2.5-7B-Instruct. Across 2 hardware configs measured on both, Qwen2.5-Coder-7B-Instruct is faster on 1 of 2. Every cell in the table below links to the submitted run it came from.

model
Qwen2.5-Coder-7B-Instruct
Qwen · 7B
View Qwen2.5-Coder-7B-Instruct page →
model
Qwen2.5-7B-Instruct
Qwen · 7B
View Qwen2.5-7B-Instruct page →

Hardware with data on both models

HardwareQwen2.5-Coder-7B-Instruct decodeQwen2.5-7B-Instruct decodeΔSource runs
RTX 5090 (32GB)255.3tok/s246.7tok/s+8.6r_2b9o6y_49mi · r_3yn-4321hp-
M3 Ultra (60-core GPU)138.6tok/s139.6tok/s-1.0r_uoehjq0nvc0 · r_5r6rhiynenc

See also: Qwen2.5-Coder-7B-Instruct benchmarks · Qwen2.5-7B-Instruct benchmarks · All hardware · All models · Methodology