Skip to content
llm-speed

Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct — every number links to the run it came from.

Verdict

Llama-3.3-70B-Instruct has no submitted run yet, so there is no head-to-head number to report. We do have 1 hardware config measured for Qwen2.5-72B-Instruct (peak 16 tok/s decode on M3 Ultra (60-core GPU)) — the table below is that baseline. Run the llm-speed suite on Llama-3.3-70B-Instruct and this page fills in the second column automatically.

model
Qwen2.5-72B-Instruct
Qwen · 72B
View Qwen2.5-72B-Instruct page →
model
Llama-3.3-70B-Instruct
Meta · 70B
View Llama-3.3-70B-Instruct page →

Hardware with data on both models

HardwareQwen2.5-72B-Instruct decodeLlama-3.3-70B-Instruct decodeΔSource runs
M3 Ultra (60-core GPU)16.31tok/s16.78tok/s-0.5r_5c80gthqlh6 · r_sx3a4y9n-m4

See also: Qwen2.5-72B-Instruct benchmarks · Llama-3.3-70B-Instruct benchmarks · All hardware · All models · Methodology