llama3.2
4 workload results across 1 hardware configuration.
Fastest local config
103.1 decode tok/s
on M1 Pro (16-core GPU) + 16GB unified via ollama (Q8_0). see full run
Local runs (4 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M1 Pro (16-core GPU) + 16GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.32.14 | Q8_0 | 103.1tok/s | 8.77tok/s | 14,247ms | r_w-p7v--eync |
| chat-long | ollama@0.32.14 | Q8_0 | 81.35tok/s | 1,246.5tok/s | 2,529ms | r_w-p7v--eync |
| concurrent-decode | ollama@0.32.14 | Q8_0 | 97.26tok/s | no data | no data | r_w-p7v--eync |
| agent-trace | ollama@0.32.14 | Q8_0 | 93.14tok/s | 2,975.3tok/s | 516ms | r_w-p7v--eync |