Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct
A single shareable card for the Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct matchup. Numbers are individual submitted decode results; each available side shows its source configuration. These are not repeated controlled comparisons.
Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct
Individual submitted decode results; hardware and workload shown below.
72B
16.3tok/s
decode (individual submitted result)
- Model
- mlx-community/qwen2.5-72b-Instruct-4bit
- Reported hardware
- M3 Ultra (60-core GPU) + 96GB unified
- Workload / suite
- chat-short · suite-v1
- Runtime
- mlx · 0.31.3
- Quantization field
- Not reported
- Submitted
- 2026-04-28
70B
16.8tok/s
decode (individual submitted result)
- Model
- mlx-community/llama-3.3-70b-Instruct-4bit
- Reported hardware
- M3 Ultra (60-core GPU) + 96GB unified
- Workload / suite
- chat-short · suite-v1
- Runtime
- mlx · 0.31.3
- Quantization field
- Not reported
- Submitted
- 2026-04-28
These independently selected results can use different models, workloads and settings. They do not establish a hardware or model winner.
Need the long-form table? Open the Qwen2.5-72B-Instruct vs Llama-3.3-70B-Instruct comparison for every overlapping (model × hardware) row, source runs, and methodology.