Skip to content
llm-speed

M4 Max (40-core GPU) LLM benchmark

The fastest LLM measured on the M4 Max (40-core GPU) is qwen3-coder-bench-32k at 113.3 decode tok/s via ollama (signed run). Across 17 reproducible runs on 2 models, this page lists decode tok/s, prefill, and TTFT for each, every number linking to the run it came from.

Fastest known config on M4 Max (40-core GPU)

113.3 decode tok/s

qwen3-coder-bench-32k via ollama (Q4_K_M). see full run

qwen3-coder-bench-32k

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortollama@0.30.11Q4_K_M113.3tok/s84.12tok/s1,308msr_roktphpc--8
chat-longollama@0.30.11Q4_K_M97.40tok/s1,342.8tok/s2,344msr_roktphpc--8
concurrent-decodeollama@0.30.11Q4_K_M109.7tok/sno datano datar_roktphpc--8
agent-traceollama@0.30.11Q4_K_M103.2tok/s3,371.6tok/s477msr_roktphpc--8

qwen3-coder

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortollama@0.30.11Q4_K_M94.51tok/s65.06tok/s1,691msr_r0di2hkku1h
chat-longollama@0.30.11Q4_K_M97.35tok/s1,397.3tok/s2,252msr_r0di2hkku1h
concurrent-decodeollama@0.30.11Q4_K_M109.9tok/sno datano datar_r0di2hkku1h
agent-traceollama@0.30.11Q4_K_Mno datano datano datar_r0di2hkku1h
chat-shortollama@0.30.11Q4_K_M55.45tok/s675.2tok/s163msr_1o8q4lhgj88
chat-longollama@0.30.11Q4_K_M57.86tok/s28,422.4tok/s111msr_1o8q4lhgj88
concurrent-decodeollama@0.30.11Q4_K_M67.01tok/sno datano datar_1o8q4lhgj88
agent-traceollama@0.30.11Q4_K_Mno datano datano datar_1o8q4lhgj88
chat-shortollama@0.30.11Q4_K_M92.70tok/s719.0tok/s153msr_40h1dznuznk
chat-longollama@0.30.11Q4_K_M94.61tok/s1,396.5tok/s2,253msr_40h1dznuznk
concurrent-decodeollama@0.30.11Q4_K_M109.1tok/sno datano datar_40h1dznuznk
agent-traceollama@0.30.11Q4_K_Mno datano datano datar_40h1dznuznk
chat-shortollama@0.30.11Q4_K_M93.17tok/s13.30tok/s8,270msr_txvkjgoyyxs

Community folklore on M4 Max (40-core GPU)

72 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.

  • communityconfidence 75%

    22.06tok/s Phi4 on M4 Max via ollama fp16

    our signed data: M4 Max · Phi4

    y **Results (All models downloaded from Ollama)** **gemma3:27b** |Quantization|Load Duration|Inference Speed| |:-|:-|:-| |q4|52.482042ms|22.06 tokens/s| |fp16|56.4445ms|6.99 tokens/s| **gemma3:12b** |Quantization|Load Duration|Inference Speed| |:-|:-|:-| |q4|56.818334ms|43.8…

    source: Reddit · u/purealgo · 2025-03-12

  • communityconfidence 70%

    25.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 70%

    17.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 70%

    25.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 70%

    17.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 70%

    25.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    ants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a surprisingly great result when placed next to ~25 tokens/s for the M4 Max. Edit: Especially when it could go even faster if only LM Studio had n-cpu-moe to put some of the load back onto …

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 70%

    17.00tok/s GPT-OSS 120B on M4 Max via lm-studio

    our signed data: M4 Max · GPT-OSS 120B

    Can you share the prompt to generate the numbers in the graph? I'm getting ~15-17 token/s with 64k context GPT-OSS 120B MXFP4 (no KV quants) in LM Studio on an i7 13800H P1G6 with 128GB DDR5-5600 and an RTX 2000 Ada, which is a s

    source: Reddit · u/Dexamph · 2025-08-18

  • communityconfidence 60%

    101.9tok/s Qwen2.5-7B on M4 Max via mlx

    our signed data: M4 Max · Qwen2.5-7B

    GPU) | | --------------------------- | -------------------------------- | ------------------------------ | | Qwen2.5-7B-Instruct (4bit) | 101.87 tokens/s | 38.99 tokens/s | | Qwen2.5-14B-Instruct (4bit) | 52.22 tokens/s | 18.88 …

    source: Reddit · u/purealgo · 2025-02-28

  • communityconfidence 60%

    8.76tok/s Qwen2.5 on M4 Max via lm-studio

    our signed data: M4 Max · Qwen2.5

    okens/s | | Qwen2.5:32B (4bit) | 19.35 tokens/s | 6.95 tokens/s | | Qwen2.5:72B (4bit) | 8.76 tokens/s | Didn't Test | #### LM Studio | MLX models | M4 Max (128 GB RAM, 40-…

    source: Reddit · u/purealgo · 2025-02-28

  • communityconfidence 60%

    6.95tok/s Qwen2.5 on M4 Max via lm-studio

    our signed data: M4 Max · Qwen2.5

    :14B (4bit) | 38.23 tokens/s | 14.66 tokens/s | | Qwen2.5:32B (4bit) | 19.35 tokens/s | 6.95 tokens/s | | Qwen2.5:72B (4bit) | 8.76 tokens/s | Didn't Test | #### LM Studio …

    source: Reddit · u/purealgo · 2025-02-28

See all 72 claims for M4 Max (40-core GPU)

Models measured on M4 Max (40-core GPU)

Common questions about M4 Max (40-core GPU)

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the M4 Max (40-core GPU) FAQ →