Skip to content
llm-speed

Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct — every number links to the run it came from.

Verdict

On the RTX 5090 (32GB), Qwen2.5-7B-Instruct decodes at 247 tok/s versus 232 tok/s for Llama-3.1-8B-Instruct, 1.1× faster. Across 3 hardware configs measured on both, Qwen2.5-7B-Instruct is faster on 3 of 3. Every cell in the table below links to the submitted run it came from.

model
Qwen2.5-7B-Instruct
Qwen · 7B
View Qwen2.5-7B-Instruct page →
model
Llama-3.1-8B-Instruct
Meta · 8B
View Llama-3.1-8B-Instruct page →

Hardware with data on both models

HardwareQwen2.5-7B-Instruct decodeLlama-3.1-8B-Instruct decodeΔSource runs
RTX 5090 (32GB)246.7tok/s232.2tok/s+14.4r_3yn-4321hp- · r_kfrkg-vn376
M3 Ultra (60-core GPU)139.6tok/s130.2tok/s+9.4r_5r6rhiynenc · r_v2pbc0rq2l4
M3 Pro (18-core GPU)30.52tok/s29.20tok/s+1.3r_llzv_g-ymaf · r_h0-use1ypnb

See also: Qwen2.5-7B-Instruct benchmarks · Llama-3.1-8B-Instruct benchmarks · All hardware · All models · Methodology