Qwen2.5-7B-Instruct-4bit
3 workload results across 2 hardware configurations.
Fastest local config
139.6 decode tok/s
on M3 Ultra (60-core GPU) + 96GB unified via mlx (4bit). see full run
Local runs (3 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M3 Pro (18-core GPU) + 36GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 30.52tok/s | 161.9tok/s | 809ms | r_llzv_g-ymaf |
| chat-short | mlx@0.31.3 | 4bit | 15.67tok/s | 54.34tok/s | 2,411ms | r_e3t93rscswq |
M3 Ultra (60-core GPU) + 96GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 139.6tok/s | 190.0tok/s | 689ms | r_5r6rhiynenc |