Skip to content
llm-speed

Fastest small-MoE coder on a Mac (Qwen3-Coder-30B, DeepSeek-Coder-V2-Lite)

Two MoE coders — Qwen3-Coder-30B-A3B and DeepSeek-Coder-V2-Lite — both clear 100 tok/s decode on an M3 Ultra at 4-bit. Here's how they rank across every Apple Silicon tier we have data for.

Verdict

As of today, the fastest small-MoE coder on a Mac on the leaderboard is DeepSeek-Coder-V2-Lite-Instruct on M3 Ultra (60-core GPU) at 168.3tok/s. Runner-up: Qwen3-Coder-30B-A3B-Instruct on M3 Ultra (60-core GPU) at 112.2tok/s. The MoE-A3B activation pattern is what makes this work on a bandwidth-bound architecture: only the active subset has to stream per token, so an Ultra with ~800 GB/s unified bandwidth keeps up.

Recommendation

Fastest small-MoE coder on Apple Silicon: DeepSeek-Coder-V2-Lite-Instruct on M3 Ultra (60-core GPU) at 168.3tok/s.

MoE coders with ~3 B active parameters are the sweet spot for Apple Silicon: the activation pattern means the bandwidth bottleneck only hits the active subset, so an M-series Ultra punches well above its raw memory-bandwidth number. Qwen3-Coder-30B-A3B and DeepSeek-Coder-V2-Lite (16B-A2.4B) both fit comfortably at 4-bit MLX on a 64 GB+ unified-memory Mac and decode fast enough for a daily-driver coding agent. Below is every submitted run we have for either model, ranked by decode tok/s, with the Mac tier (Pro / Max / Ultra) called out so the price-to-tok/s line is legible. If your config isn't there yet, run the suite and submit — that row will appear next refresh.

Submitted benchmarks

Side-by-side comparisons

See also: All hardware · All models · Methodology