Fastest small-MoE coder on a Mac (Qwen3-Coder-30B, DeepSeek-Coder-V2-Lite)
Two MoE coders — Qwen3-Coder-30B-A3B and DeepSeek-Coder-V2-Lite — both clear 100 tok/s decode on an M3 Ultra at 4-bit. Here's how they rank across every Apple Silicon tier we have data for.
As of today, the fastest small-MoE coder on a Mac on the leaderboard is DeepSeek-Coder-V2-Lite-Instruct on M3 Ultra (60-core GPU) at 168.3tok/s. Runner-up: Qwen3-Coder-30B-A3B-Instruct on M3 Ultra (60-core GPU) at 112.2tok/s. The MoE-A3B activation pattern is what makes this work on a bandwidth-bound architecture: only the active subset has to stream per token, so an Ultra with ~800 GB/s unified bandwidth keeps up.
Fastest small-MoE coder on Apple Silicon: DeepSeek-Coder-V2-Lite-Instruct on M3 Ultra (60-core GPU) at 168.3tok/s.
MoE coders with ~3 B active parameters are the sweet spot for Apple Silicon: the activation pattern means the bandwidth bottleneck only hits the active subset, so an M-series Ultra punches well above its raw memory-bandwidth number. Qwen3-Coder-30B-A3B and DeepSeek-Coder-V2-Lite (16B-A2.4B) both fit comfortably at 4-bit MLX on a 64 GB+ unified-memory Mac and decode fast enough for a daily-driver coding agent. Below is every submitted run we have for either model, ranked by decode tok/s, with the Mac tier (Pro / Max / Ultra) called out so the price-to-tok/s line is legible. If your config isn't there yet, run the suite and submit — that row will appear next refresh.
Submitted benchmarks
| Mac | Model | decode tok/s | Workload | Run |
|---|---|---|---|---|
| M3 Ultra (60-core GPU) | DeepSeek-Coder-V2-Lite-Instruct | 168.3tok/s | chat-short | r_l_v1-zq_qaz |
| M3 Ultra (60-core GPU) | Qwen3-Coder-30B-A3B-Instruct | 112.2tok/s | chat-short | r_fpsca03u2o_ |
Side-by-side comparisons
See also: All hardware · All models · Methodology