Skip to content
llm-speed
Leaderboard/models/qwen3-4b-instruct-2507

Qwen3-4B-Instruct-2507-4bit

4 workload results across 1 hardware configuration.

Fastest local config

182.7 decode tok/s

on M3 Ultra (60-core GPU) + 96GB unified via mlx. see full run

Local runs (4 runs)

Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.

M3 Ultra (60-core GPU) + 96GB unifiedM3 Ultra (60-core GPU) + 96GB unified

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortmlx@0.31.3-182.7tok/s636.6tok/s173msr_p2rpb_0iyjj
chat-longmlx@0.31.3-147.0tok/s1,913.7tok/s1,644msr_p2rpb_0iyjj
concurrent-decodemlx@0.31.3-168.4tok/sno datano datar_p2rpb_0iyjj
agent-tracemlx@0.31.3-160.8tok/s1,913.6tok/s1,091msr_p2rpb_0iyjj

Qwen3-4B-Instruct-2507-4bit on hardware