Skip to content
llm-speed

DeepSeek-Coder-V2-Lite-Instruct vs Qwen3-Coder-30B-A3B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for DeepSeek-Coder-V2-Lite-Instruct vs Qwen3-Coder-30B-A3B-Instruct — every number links to the run it came from.

How to read these results

Selected submissions for RTX 5090 (32GB) report 309 tok/s on DeepSeek-Coder-V2-Lite-Instruct and 260 tok/s on Qwen3-Coder-30B-A3B-Instruct. The table selects each side's fastest submitted row independently, grouped by hardware label. Matching labels do not establish matching model artifacts, quantization, backend, workload or context. These peaks are reference results, not a controlled speed or answer-quality ranking. Inspect both source runs before choosing hardware.

model
DeepSeek-Coder-V2-Lite-Instruct
DeepSeek · 16B-A2.4B
View DeepSeek-Coder-V2-Lite-Instruct page →
model
Qwen3-Coder-30B-A3B-Instruct
Qwen · 30B-A3B
View Qwen3-Coder-30B-A3B-Instruct page →

Hardware with data on both models

HardwareDeepSeek-Coder-V2-Lite-Instruct decodeQwen3-Coder-30B-A3B-Instruct decodeΔSource runs
RTX 5090 (32GB)309.5tok/s
chat-short · llama.cpp · quantization unreported
259.9tok/s
chat-short · llama.cpp · quantization unreported
+49.6r_0_gs1rgl2fl · r_c7qyvvmmsv1
M3 Ultra (60-core GPU)168.3tok/s
chat-short · mlx · quantization unreported
112.2tok/s
chat-short · mlx · quantization unreported
+56.2r_l_v1-zq_qaz · r_fpsca03u2o_

See also: DeepSeek-Coder-V2-Lite-Instruct benchmarks · Qwen3-Coder-30B-A3B-Instruct benchmarks · All hardware · All models · Methodology