Skip to content
llm-speed
Leaderboard/models/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct-4bit

6 workload results across 2 hardware configurations.

Fastest local config

259.9 decode tok/s

on RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB via llama.cpp. see full run

Local runs (6 workload results)

Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.

RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBRTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortllama.cpp-68.41tok/sno data334msr_pm_a1uf2ufc
chat-longllama.cpp-70.55tok/sno data1,204msr_pm_a1uf2ufc
concurrent-decodellama.cpp-68.04tok/sno datano datar_pm_a1uf2ufc
agent-tracellama.cpp-66.09tok/s4,816.7tok/s451msr_pm_a1uf2ufc
chat-shortllama.cpp-259.9tok/sno data218msr_c7qyvvmmsv1

M3 Ultra (60-core GPU) + 96GB unifiedM3 Ultra (60-core GPU) + 96GB unified

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortmlx@0.31.3-112.2tok/s204.0tok/s539msr_fpsca03u2o_

Qwen3-Coder-30B-A3B-Instruct-4bit on hardware