Skip to content
llm-speed

M3 Ultra vs M4 Max

A single shareable card for the M3 Ultra vs M4 Max matchup. Numbers are individual submitted decode results; each available side shows its source configuration. These are not repeated controlled comparisons.

M3 Ultra vs M4 Max

Individual submitted decode results; model and workload shown below.

Apple M3
429.2tok/s
decode (individual submitted result)
Model
mlx-community/Qwen3-0.6B-4bit
Reported hardware
M3 Ultra (60-core GPU) + 96GB unified
Workload / suite
chat-short · suite-v1
Runtime
mlx · 0.31.3
Quantization field
Not reported
Submitted
2026-07-12
source: r_6rgao6ao735
Apple M4
113.3tok/s
decode (individual submitted result)
Model
qwen3-coder-bench-32k
Reported hardware
M4 Max (40-core GPU) + 128GB unified
Workload / suite
chat-short · suite-v1
Runtime
ollama · 0.30.11
Quantization field
Q4_K_M
Submitted
2026-06-27
source: r_roktphpc--8
These independently selected results can use different models, workloads and settings. They do not establish a hardware or model winner.

The memory shown in a source configuration belongs to that submission. For M3 Ultra purchases, Apple’s UK specification checked September 6, 2026 lists 96 GB configurable to 256 GB. Its 2025 launch announcement described up to 512 GB. Confirm the exact configuration; a family maximum is not the machine measured here.

Need the long-form table? Open the M3 Ultra vs M4 Max comparison for every overlapping (model × hardware) row, source runs, and methodology.