Skip to content
llm-speed

Cheatsheet

Local LLM speed cheatsheet: decode tok/s by GPU & Mac

Best decode tokens-per-second per (model × hardware) tuple measured by llm-speed under suite-v1. Numbers are wall-clock, batch size 1 unless the workload says otherwise. Each row links to the canonical run page; cite as llm-speed.com/r/<id>.

36 (model × hardware) combinations · sorted by decode tok/s · suite-v1

ModelQuantHardwareBackendWorkloaddecode tok/sprefill tok/sTTFT (ms)Run
mlx-community/Qwen3-0.6B-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short429.21185.292.8r_6rgao6ao735
stable-code-instruct-3b-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode356.1r_q9f15lz6831
gpt-oss-20b-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short318.4172.5r_b9ul-vxh9sc
DeepSeek-Coder-V2-Lite-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short293.1297.0r_bfpto9so2o1
Qwen2.5-Coder-7B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short255.335.4r_2b9o6y_49mi
Qwen2.5-7B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode240.7r_1shiviswt3d
Llama-3.1-8B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short229.435.6r_qr4srge34da
gemma3Q4_K_MRTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GBollama@0.31.1chat-short195.0167.2705.6r_dlanfbgym0h
deepseek-coder-v2Q4_0RTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GBollama@0.31.1chat-short189.5159.2734.9r_o2-1w665rtq
mlx-community/Qwen3-4B-Instruct-2507-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short182.7636.6172.8r_p2rpb_0iyjj
qwen3-coderQ4_K_MRTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GBollama@0.31.1concurrent-decode179.9r_wjq32z47vlp
qwen2.5-coderQ4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1concurrent-decode161.1r_mv8n8k9wu1e
llama3.1Q4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1chat-short154.4328.0335.4r_h1ub_1uxzdh
gemma-4-12b-it-qatQ4_0RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cpp@1 (9725a31)long-context-decay142.67115.83354.8r_-lh-vczf7ik
gpt-ossMXFP4RTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GBollama@0.31.1chat-long141.7631.55048.2r_iu2sfa9ykvw
phi-4-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode140.2r_k1u_k1j_1i2
Qwen2.5-Coder-14B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short140.042.1r_opuj21f13-_
qwen2.5-coderQ4_K_MRTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GBollama@0.31.1concurrent-decode139.2r_pb0jpnbujji
llama3.1Q4_K_MRTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GBollama@0.31.1chat-short136.22.249960.0r_9hurqggbshk
deepseek-r1Q4_K_MRTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GBollama@0.31.1agent-trace133.88852.4366.2r_fg77v2hhohb
Qwen2.5-14B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode131.4r_xr4qdv1hgf2
glm-4.7-flashQ4_K_MRTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GBollama@0.31.1agent-trace129.92491.11262.5r_o636l3cc-rr
mlx-community/Qwen3-8B-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short120.5446.1246.6r_b1awwhkoo2v
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short116.8424.5247.4r_yy_6jfx70jq
mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short112.5308.7356.3r_o6wbmfegmde
llama3.2Q8_0M1 Pro (16-core GPU) + 16GB unifiedollama@0.32.14chat-short103.18.814246.7r_w-p7v--eync
qwen3.6-agent-q6k-ctx128kQ6_KM5 Max (40-core GPU) + 48GB unifiedollama@0.32.1chat-short88.316.26915.3r_udwgba7udqu
Yi-Coder-9B-Chat-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppagent-trace71.84478.0397.2r_zlh6az5q0o_
Qwen2.5-Coder-32B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode71.0r_v983y0y3r2u
Qwen3-Coder-30B-A3B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-long70.61203.7r_pm_a1uf2ufc
gemma-2-9b-it-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short69.5324.6r_1_xl4zb5-xj
qwen2.5-coderQ4_K_MRTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GBollama@0.31.1chat-short69.2402.0325.8r_6ahy-dq0f_0
Qwen2.5-32B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short68.777.0r_bjy5a5izxjc
Qwen3-32B-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode66.6r_-txe_hiq44n
Qwen3.8-27BQ4_K_MRTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cpp@1 (9725a31)chat-short65.7973.6117.1r_c3ps9wjygi9
qwen3.6Q4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1agent-trace44.21910.51807.9r_h_659oy695r

How to read this table

Each row is the highest decode tokens-per-second measured for a unique (model, hardware) pair. When more than one workload was run, we pick the workload that produced the headline number. Lower-tps duplicates from the same machine are not shown; click through to the run page for the full per-workload breakdown.

decode tok/s is wall-clock streaming throughput (memory-bandwidth-bound). prefill tok/s is prompt ingestion throughput (compute-bound). TTFT is time-to-first-token in milliseconds. See /glossary for definitions and /methodology for the workload spec.