Skip to content
llm-speed

M4 Pro LLM benchmark

No M4 Pro LLM benchmarks yet. Be the first to submit a signed, reproducible tok/s run from your own M4 Pro.

Will it run local LLMs?

No signed M4 Pro run yet, but its 64GB of unified memory tells you what fits. At 4-bit it comfortably runs models up to about 57B parameters with room for context, and up to roughly 91B with tighter quantization. Check a specific model with the VRAM-fit tool.

For measured decode speed on the nearest hardware we have benchmarked: M3 Ultra · M4 Max · M3 Pro. Run the suite on your M4 Pro to fill this page in.

No M4 Pro benchmarks yet.

Run on YOUR hardware to populate this page: pipx install llm-speed && llm-speed bench

$ pipx install llm-speed && llm-speed bench

Community folklore on M4 Pro

21 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.

  • communityconfidence 70%

    70.00tok/s Qwen3-Coder on M4 Pro via mlx

    our signed data: M4 Pro · Qwen3-Coder

    urpose-built for code, MoE architecture so only 3B params active per token. Fits in 36GB with room for 16-32K context. On M4 Pro MLX I get ~70 tok/s with it. If you also want a general-purpose model to keep alongside it, **Qwen3.5-35B-A3B** is the same MoE architecture, similar …

    source: Reddit · u/the_real_druide67 · 2026-03-28

  • communityconfidence 60%

    7.24tok/s qwen2.5 on M4 Pro via ollama

    our signed data: M4 Pro · qwen2.5

    2.5 right up until you hit the 72B parameter size. Due to VRAM it falls apart. The M4 max and M4 Pro could run the 72B, but at 8.8t/s and 7.24t/s. This is certainly better than running on a CPU (which is what happens with the 4090), but it's still too slow to be worth a $4k pu…

    source: Reddit · u/darth_chewbacca · 2024-11-14

  • communityconfidence 60%

    8.80tok/s qwen2.5 on M4 Pro via ollama

    our signed data: M4 Pro · qwen2.5

    lly at qwen2.5 right up until you hit the 72B parameter size. Due to VRAM it falls apart. The M4 max and M4 Pro could run the 72B, but at 8.8t/s and 7.24t/s. This is certainly better than running on a CPU (which is what happens with the 4090), but it's still too slow to be wor…

    source: Reddit · u/darth_chewbacca · 2024-11-14

  • communityconfidence 60%

    72.00tok/s GPT OSS 20B on M4 Pro via mlx

    our signed data: M4 Pro · GPT OSS 20B

    edge frozen as of 2024-06, handled rescaling and converting a recipe into metric that Qwen 3 30B 3AB 2507 completely hosed, churns at about 72 tok/s on a MacBook M4 Pro Max 4 bit MLX. And it ALMOST got a side scrolling shooter working, whereas the 120B didn't, and Qwen 3 didn't …

    source: Reddit · u/cspenn · 2025-08-05

  • communityconfidence 60%

    14.50tok/s Gemma 3 27b on M4 Pro via mlx

    our signed data: M4 Pro · Gemma 3 27b

    Mac Mini M4 Pro (20c GPU) 64GB unified RAM; Gemma 3 27b with MLX 14.5t/s power usage almost 70W (including connected keyboard and mouse). So more efficient then even the Ryzen 395 AI (if those results are accurat

    source: Reddit · u/Cergorach · 2025-07-11

  • communityconfidence 60%

    55.00tok/s Qwen3 30b a3b on M4 Pro via mlx

    our signed data: M4 Pro · Qwen3 30b a3b

    Binned M4 Pro/48GB owner here since November--current daily driver is Qwen3 30b a3b 8-bit MLX @ 55t/s ymmv, but I like it a lot and it flies.

    source: Reddit · u/MrPecunius · 2025-07-23

  • communityconfidence 60%

    11.37tok/s Qwen3 30b-a3b on M4 Pro via mlx

    our signed data: M4 Pro · Qwen3 30b-a3b

    That's about three times as fast as my binned M4 Pro/48GB: with Qwen3 30b-a3b 8-bit MLX, I got 180 seconds to first token and 11.37t/s with the same size lorem ipsum prompt. That tracks really well with the 3X memory bandwidth difference.

    source: Reddit · u/MrPecunius · 2025-05-26

  • communityconfidence 55%

    60.50tok/s on M4 Pro via lm-studio FP16

    our signed data: M4 Pro

    32 GB DDR5|CUDA 12 llama.cpp (LM Studio)|59.1 tok/s|0.02 s|\-| |M4 Pro|16 GPU cores, MacBook Pro 14”, 48 GB unified memory|MLX (LM Studio)|60.5 tok/s 👑|0.31 s|3.69| # Super Interesting Notes: **1. The neural accelerators didn't make much of a difference. Here's why!** * First …

    source: Reddit · u/TechExpert2910 · 2025-10-27

  • communityconfidence 55%

    60.50tok/s on M4 Pro via lm-studio FP16

    our signed data: M4 Pro

    32 GB DDR5|CUDA 12 llama.cpp (LM Studio)|59.1 tok/s|0.02 s|\-| |M4 Pro|16 GPU cores, MacBook Pro 14”, 48 GB unified memory|MLX (LM Studio)|60.5 tok/s 👑|0.31 s|3.69| # Super Interesting Notes: **1. The neural accelerators didn't make much of a difference. Here's why!** * First …

    source: Reddit · u/TechExpert2910 · 2025-10-27

  • communityconfidence 55%

    60.50tok/s on M4 Pro via lm-studio FP16

    our signed data: M4 Pro

    32 GB DDR5|CUDA 12 llama.cpp (LM Studio)|59.1 tok/s|0.02 s|\-| |M4 Pro|16 GPU cores, MacBook Pro 14”, 48 GB unified memory|MLX (LM Studio)|60.5 tok/s 👑|0.31 s|3.69| # Super Interesting Notes: **1. The neural accelerators didn't make much of a difference. Here's why!** * First …

    source: Reddit · u/TechExpert2910 · 2025-10-27

See all 21 claims for M4 Pro

Common questions about M4 Pro

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the M4 Pro FAQ →