Qwen3-8B-4bit
4 workload results across 1 hardware configuration.
Fastest local config
120.5 decode tok/s
on M3 Ultra (60-core GPU) + 96GB unified via mlx. see full run
Local runs (4 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
M3 Ultra (60-core GPU) + 96GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 120.5tok/s | 446.1tok/s | 247ms | r_b1awwhkoo2v |
| chat-long | mlx@0.31.3 | - | 103.3tok/s | 1,108.8tok/s | 2,838ms | r_b1awwhkoo2v |
| concurrent-decode | mlx@0.31.3 | - | 114.3tok/s | no data | no data | r_b1awwhkoo2v |
| agent-trace | mlx@0.31.3 | - | 109.4tok/s | 1,109.1tok/s | 1,885ms | r_b1awwhkoo2v |
Community folklore
9 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 75%
92.40tok/s — Qwen3-8B on RTX 4090 via llama.cpp Q8_0
our signed data: RTX 4090 · Qwen3-8B
“tok/s` generation at `128` output tokens - `Q8_0`: about `9975 tok/s` prompt processing at `512` tokens, `9955 tok/s` at `1024`, and about `92.4 tok/s` generation at `128` output tokens Hardware / runtime for those numbers: - `RTX 4090` - `Ryzen 9 7900X` - `llama.cpp` build com…”
- communityconfidence 75%
92.40tok/s — Qwen3-8B on RTX 4090 via llama.cpp Q8_0
our signed data: RTX 4090 · Qwen3-8B
“tok/s` generation at `128` output tokens - `Q8_0`: about `9975 tok/s` prompt processing at `512` tokens, `9955 tok/s` at `1024`, and about `92.4 tok/s` generation at `128` output tokens Hardware / runtime for those numbers: - `RTX 4090` - `Ryzen 9 7900X` - `llama.cpp` build com…”
- communityconfidence 75%
79.60tok/s — Qwen3-8B on M3 Max via mlx bf16
our signed data: M3 Max · Qwen3-8B
“Happy to coordinate if you open-source yours first — no point duplicating. EDIT: Benchmarks complete on M3 Max (128GB). I am getting 79.6 tok/s on Qwen3-8B-bf16 (3.41x speedup) with confirmed bit-for-bit parity against the baseline. For those interested in the MLX-specific …”
- communityconfidence 75%
79.60tok/s — Qwen3-8B on M3 Max via mlx bf16
our signed data: M3 Max · Qwen3-8B
“Happy to coordinate if you open-source yours first — no point duplicating. EDIT: Benchmarks complete on M3 Max (128GB). I am getting 79.6 tok/s on Qwen3-8B-bf16 (3.41x speedup) with confirmed bit-for-bit parity against the baseline. For those interested in the MLX-specific …”
- communityconfidence 75%
79.60tok/s — Qwen3-8B on M3 Max via mlx bf16
our signed data: M3 Max · Qwen3-8B
“Happy to coordinate if you open-source yours first — no point duplicating. EDIT: Benchmarks complete on M3 Max (128GB). I am getting 79.6 tok/s on Qwen3-8B-bf16 (3.41x speedup) with confirmed bit-for-bit parity against the baseline. For those interested in the MLX-specific …”
- communityconfidence 75%
79.60tok/s — Qwen3-8B on M3 Max via mlx bf16
our signed data: M3 Max · Qwen3-8B
“Happy to coordinate if you open-source yours first — no point duplicating. EDIT: Benchmarks complete on M3 Max (128GB). I am getting 79.6 tok/s on Qwen3-8B-bf16 (3.41x speedup) with confirmed bit-for-bit parity against the baseline. For those interested in the MLX-specific …”
- communityconfidence 60%
73.00tok/s — Qwen3 8B on RTX 5060 Ti via llama.cpp
our signed data: RTX 5060 Ti · Qwen3 8B
“e even pytorch via mps???). A completely useless graph. I can also provide the same useless information: the Qwen3 8B model on my M1 gives 73 t/s and time to first token = 0.65s, which is faster than your rtx 5060ti. Now guess what type of M1 I have, what runtime I used, and wha…”
- communityconfidence 55%
60.00tok/s — Qwen3-8B via vllm fp4
our signed data: Qwen3-8B
“ance. So benchmarked two models for local inference: 1. Ollama serving qwen3:8b-q4\_K\_M = 70 t/s 2. VLLM serving nvidia/Qwen3-8B-NVFP4 = 60 t/s Both generated \~1000 tokens on a simple 50-token prompt. The token generation performance was reported via \`--verbose\` flag in ol…”
- communityconfidence 45%
711.0tok/s — Qwen3-8B Q8_0
our signed data: Qwen3-8B
“Run with --flash-attn ``` **Dense model (Qwen3-8B Q8_0) — prompt processing:** - ROCm default, no flash attn: **711 t/s** - ROCm + flash attn only: **~3,980 t/s** - **5.5× improvement from one flag** --- ## …”