Qwen2.5-0.5B-Instruct-4bit
2 workload results across 1 hardware configuration.
Fastest local config
286.5 decode tok/s
on M3 Pro (18-core GPU) + 36GB unified via mlx (4bit). see full run
Local runs (2 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M3 Pro (18-core GPU) + 36GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | 4bit | 286.5tok/s | 656.2tok/s | 200ms | r_akcbpx5vcqa |
| chat-short | mlx@0.31.3 | 4bit | 282.6tok/s | 668.8tok/s | 196ms | r_bftqtkilvoe |