Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct
Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct — every number links to the run it came from.
Verdict
On the RTX 5090 (32GB), Qwen2.5-7B-Instruct decodes at 247 tok/s versus 232 tok/s for Llama-3.1-8B-Instruct, 1.1× faster. Across 3 hardware configs measured on both, Qwen2.5-7B-Instruct is faster on 3 of 3. Every cell in the table below links to the submitted run it came from.
Hardware with data on both models
| Hardware | Qwen2.5-7B-Instruct decode | Llama-3.1-8B-Instruct decode | Δ | Source runs |
|---|---|---|---|---|
| RTX 5090 (32GB) | 246.7tok/s | 232.2tok/s | +14.4 | r_3yn-4321hp- · r_kfrkg-vn376 |
| M3 Ultra (60-core GPU) | 139.6tok/s | 130.2tok/s | +9.4 | r_5r6rhiynenc · r_v2pbc0rq2l4 |
| M3 Pro (18-core GPU) | 30.52tok/s | 29.20tok/s | +1.3 | r_llzv_g-ymaf · r_h0-use1ypnb |
See also: Qwen2.5-7B-Instruct benchmarks · Llama-3.1-8B-Instruct benchmarks · All hardware · All models · Methodology