Skip to content
llm-speed

DGX Spark vs M3 Ultra

Signed, community-submitted decode tok/s, prefill, and TTFT for DGX Spark vs M3 Ultra — every number links to the run it came from.

Verdict

DGX Spark has no submitted run yet, so there is no head-to-head number to report. We do have 26 models measured for M3 Ultra (peak 429 tok/s decode on mlx-community/Qwen3-0.6B-4bit) — the table below is that baseline. Run the llm-speed suite on DGX Spark and this page fills in the second column automatically.

hardware
DGX Spark
NVIDIA GB10 Grace Blackwell · up to 128 GB unified memory
View DGX Spark page →
hardware
M3 Ultra
Apple M3 · up to 192 GB unified memory
View M3 Ultra page →

M3 Ultra benchmarks — no shared model with DGX Spark yet

Modeldecode tok/sWorkloadRun
mlx-community/Qwen3-0.6B-4bit429.2tok/schat-shortr_6rgao6ao735
mlx-community/stable-code-instruct-3b-4bit192.5tok/schat-shortr_y2_5y8oo97d
mlx-community/Qwen3-4B-Instruct-2507-4bit182.7tok/schat-shortr_p2rpb_0iyjj
mlx-community/DeepSeek-Coder-V2-Lite-Instruct-4bit168.3tok/schat-shortr_l_v1-zq_qaz
mlx-community/gpt-oss-20b-MXFP4-Q4152.7tok/schat-shortr_3ijun8ltjnb
mlx-community-Qwen2.5-7B-Instruct-4bit139.6tok/schat-shortr_5r6rhiynenc
mlx-community/Qwen2.5-Coder-7B-Instruct-4bit138.6tok/schat-shortr_uoehjq0nvc0
mlx-community/Llama-3.1-8B-Instruct-4bit130.2tok/schat-shortr_v2pbc0rq2l4
mlx-community/Qwen3-8B-4bit120.5tok/schat-shortr_b1awwhkoo2v
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit116.8tok/schat-shortr_yy_6jfx70jq
mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit112.5tok/schat-shortr_o6wbmfegmde
mlx-community/granite-8b-code-instruct-4bit112.3tok/schat-shortr_bue3bee0gw7
mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit112.2tok/schat-shortr_fpsca03u2o_
mlx-community/Yi-Coder-9B-Chat-4bit103.5tok/schat-shortr_3hvui9a1yuc
mlx-community/gemma-2-9b-it-4bit89.45tok/schat-shortr_iz137eqvuzy
lmstudio-community-Qwen3-Next-80B-A3B-Instruct-MLX-4bit80.34tok/schat-shortr_1pl79r50ofy
mlx-community/phi-4-4bit74.37tok/schat-shortr_sqzp-0rdez-
mlx-community-Qwen2.5-14B-Instruct-4bit70.85tok/schat-shortr_v4bq1sviz4o
mlx-community/Qwen2.5-Coder-14B-Instruct-4bit70.51tok/schat-shortr_l36cijqxq4t
mlx-community/starcoder2-15b-4bit64.16tok/schat-shortr_wsxml_39dh_
mlx-community/Codestral-22B-v0.1-4bit47.49tok/schat-shortr_79dvtag5fd_
mlx-community-Qwen2.5-32B-Instruct-4bit34.60tok/schat-shortr_njgxtgyym1e
mlx-community/Qwen2.5-Coder-32B-Instruct-4bit34.48tok/schat-shortr_721b4bls_oq
mlx-community/Qwen3-32B-4bit34.41tok/schat-shortr_anmmc80-aoq
mlx-community/llama-3.3-70b-Instruct-4bit16.78tok/schat-shortr_sx3a4y9n-m4
mlx-community/qwen2.5-72b-Instruct-4bit16.31tok/schat-shortr_5c80gthqlh6

See also: DGX Spark benchmarks · M3 Ultra benchmarks · All hardware · All models · Methodology