qwen3.6-agent-q6k-ctx128k
1 workload result across 1 hardware configuration.
Fastest local config
88.3 decode tok/s
on M5 Max (40-core GPU) + 48GB unified via ollama (Q6_K). see full run
Local runs (1 run)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M5 Max (40-core GPU) + 48GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.32.1 | Q6_K | 88.29tok/s | 16.20tok/s | 6,915ms | r_udwgba7udqu |