Skip to content
llm-speed

Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct

A single shareable card for the Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct matchup. Numbers are individual submitted decode results; each available side shows its source configuration. These are not repeated controlled comparisons.

Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct

Individual submitted decode results; hardware and workload shown below.

72B
16.3tok/s
decode (individual submitted result)
Model
mlx-community/qwen2.5-72b-Instruct-4bit
Reported hardware
M3 Ultra (60-core GPU) + 96GB unified
Workload / suite
chat-short · suite-v1
Runtime
mlx · 0.31.3
Quantization field
Not reported
Submitted
2026-04-28
source: r_5c80gthqlh6
70B
16.8tok/s
decode (individual submitted result)
Model
mlx-community/llama-3.3-70b-Instruct-4bit
Reported hardware
M3 Ultra (60-core GPU) + 96GB unified
Workload / suite
chat-short · suite-v1
Runtime
mlx · 0.31.3
Quantization field
Not reported
Submitted
2026-04-28
source: r_sx3a4y9n-m4
These independently selected results can use different models, workloads and settings. They do not establish a hardware or model winner.

Need the long-form table? Open the Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct comparison for every overlapping (model × hardware) row, source runs, and methodology.