Skip to content
llm-speed

Qwen2.5-Coder-7B-Instruct vs DeepSeek-Coder-V2-Lite-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-Coder-7B-Instruct vs DeepSeek-Coder-V2-Lite-Instruct — every number links to the run it came from.

How to read these results

Selected submissions for RTX 5090 (32GB) report 255 tok/s on Qwen2.5-Coder-7B-Instruct and 309 tok/s on DeepSeek-Coder-V2-Lite-Instruct. The table selects each side's fastest submitted row independently, grouped by hardware label. Matching labels do not establish matching model artifacts, quantization, backend, workload or context. These peaks are reference results, not a controlled speed or answer-quality ranking. Inspect both source runs before choosing hardware.

model
Qwen2.5-Coder-7B-Instruct
Qwen · 7B
View Qwen2.5-Coder-7B-Instruct page →
model
DeepSeek-Coder-V2-Lite-Instruct
DeepSeek · 16B-A2.4B
View DeepSeek-Coder-V2-Lite-Instruct page →

Hardware with data on both models

HardwareQwen2.5-Coder-7B-Instruct decodeDeepSeek-Coder-V2-Lite-Instruct decodeΔSource runs
RTX 5090 (32GB)255.3tok/s
chat-short · llama.cpp · quantization unreported
309.5tok/s
chat-short · llama.cpp · quantization unreported
-54.2r_2b9o6y_49mi · r_0_gs1rgl2fl
M3 Ultra (60-core GPU)138.6tok/s
chat-short · mlx · quantization unreported
168.3tok/s
chat-short · mlx · quantization unreported
-29.8r_uoehjq0nvc0 · r_l_v1-zq_qaz

See also: Qwen2.5-Coder-7B-Instruct benchmarks · DeepSeek-Coder-V2-Lite-Instruct benchmarks · All hardware · All models · Methodology