R1-0528-Qwen3-8B-MLX-4bit
4 workload results across 1 hardware configuration.
Fastest local config
116.8 decode tok/s
on M3 Ultra (60-core GPU) + 96GB unified via mlx. see full run
Local runs (4 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M3 Ultra (60-core GPU) + 96GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 116.8tok/s | 424.5tok/s | 247ms | r_yy_6jfx70jq |
| chat-long | mlx@0.31.3 | - | 100.5tok/s | 1,104.0tok/s | 2,846ms | r_yy_6jfx70jq |
| concurrent-decode | mlx@0.31.3 | - | 110.6tok/s | no data | no data | r_yy_6jfx70jq |
| agent-trace | mlx@0.31.3 | - | 107.2tok/s | 1,108.9tok/s | 1,884ms | r_yy_6jfx70jq |