Skip to content
llm-speed

Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct

A single shareable card for the Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct matchup. Numbers are the best decode tok/s submitted on the llm-speed suite — every side links back to the run it came from.

Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct

Best decode tok/s across every rig submitted for each model.

Qwen · 7B
247tok/s
decode (best submitted run)
source: r_3yn-4321hp-
Meta · 8B
232tok/s
decode (best submitted run)
source: r_kfrkg-vn376
Qwen2.5-7B-Instruct ← faster +14.4 tok/s (1.06× faster)

Need the long-form table? Open the Qwen2.5-7B-Instruct vs Llama-3.1-8B-Instruct comparison for every overlapping (model × hardware) row, source runs, and methodology.