Skip to content
llm-speed

qwen3.6-agent-q6k-ctx128k on M5 Max (40-core GPU) + 48GB unified

M5 Max (40-core GPU) + 48GB unifiedM5 Max (40-core GPU) + 48GB unified
suite suite-v1
cli 0.0.5
signed45bysmL5QL…
Embed badgesubmitted Jul 26, 2026

Workload results

WorkloadBackendModeldecode tok/sprefill tok/sTTFTp50p95
chat-shortollama@0.32.1qwen3.6-agent-q6k-ctx128kQ6_K88.29tok/s16.20tok/s6,915ms11.3ms11.5ms

Reproduce on your machine

Same workload, same model, signed at your rig. The exact command that produced this run:

$ pipx install llm-speed && llm-speed bench --model 'qwen3.6-agent-q6k-ctx128k' --workload 'chat-short'

Runs in about a minute. Your number lands on the leaderboard signed and linkable. How it's measured.

Embed this run

Drop the badge into a README, blog post, or signature. Each render is a backlink to the signed result.

llm-speed: 88.3 tok/s on M5 Max (40-core GPU) (qwen3.6-agent-q6k-ctx128k)
[![llm-speed: 88.3 tok/s on M5 Max (40-core GPU) (qwen3.6-agent-q6k-ctx128k)](https://llm-speed.com/badge/r_udwgba7udqu.svg)](https://llm-speed.com/r/r_udwgba7udqu)

Related benchmarks

Provenance

Run ID
r_udwgba7udqu
Fingerprint hash
2f1a9da045ba1233
Public key
45bysmL5QLTndgUhbqVKY1CgIC6kjhaOyRZ+guHHOio=
Received
2026-07-26 23:18:57