Skip to content
llm-speed

Cheatsheet

Local LLM speed cheatsheet: decode tok/s by GPU & Mac

Best decode tokens-per-second per (model × hardware) tuple measured by llm-speed under suite-v1. Numbers are wall-clock, batch size 1 unless the workload says otherwise. Each row links to the canonical run page; cite as llm-speed.com/r/<id>.

33 (model × hardware) combinations · sorted by decode tok/s · suite-v1

ModelQuantHardwareBackendWorkloaddecode tok/sprefill tok/sTTFT (ms)Run
mlx-community/Qwen3-0.6B-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short429.21185.292.8r_6rgao6ao735
stable-code-instruct-3b-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode356.1r_q9f15lz6831
gpt-oss-20b-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short318.4172.5r_b9ul-vxh9sc
DeepSeek-Coder-V2-Lite-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short309.5267.6r_0_gs1rgl2fl
Qwen3-Coder-30B-A3B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short259.9218.4r_c7qyvvmmsv1
Qwen2.5-Coder-7B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short255.335.4r_2b9o6y_49mi
Qwen2.5-7B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short246.733.9r_3yn-4321hp-
Llama-3.1-8B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short232.233.8r_kfrkg-vn376
Yi-Coder-9B-Chat-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short199.336.2r_u4iojm6-ekg
gemma3Q4_K_MRTX 4090 (24GB) + AMD EPYC 7443 24-Core Processor (24c) + 252GBollama@0.31.1chat-short195.0167.2705.6r_dlanfbgym0h
deepseek-coder-v2Q4_0RTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GBollama@0.31.1chat-short189.5159.2734.9r_o2-1w665rtq
mlx-community/Qwen3-4B-Instruct-2507-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short182.7636.6172.8r_p2rpb_0iyjj
qwen3-coderQ4_K_MRTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GBollama@0.31.1concurrent-decode179.9r_wjq32z47vlp
qwen2.5-coderQ4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1concurrent-decode161.1r_mv8n8k9wu1e
llama3.1Q4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1chat-short154.4328.0335.4r_h1ub_1uxzdh
gpt-ossMXFP4RTX 4090 (24GB) + AMD EPYC 7352 24-Core Processor (24c) + 252GBollama@0.31.1chat-long141.7631.55048.2r_iu2sfa9ykvw
phi-4-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode140.2r_k1u_k1j_1i2
Qwen2.5-Coder-14B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short140.042.1r_opuj21f13-_
qwen2.5-coderQ4_K_MRTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GBollama@0.31.1concurrent-decode139.2r_pb0jpnbujji
llama3.1Q4_K_MRTX 3090 (24GB) + AMD EPYC 7663 56-Core Processor (56c) + 252GBollama@0.31.1chat-short136.22.249960.0r_9hurqggbshk
deepseek-r1Q4_K_MRTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GBollama@0.31.1agent-trace133.88852.4366.2r_fg77v2hhohb
Qwen2.5-14B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short133.355.9r_tj9bu7gvnvh
glm-4.7-flashQ4_K_MRTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GBollama@0.31.1agent-trace129.92491.11262.5r_o636l3cc-rr
mlx-community/Qwen3-8B-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short120.5446.1246.6r_b1awwhkoo2v
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short116.8424.5247.4r_yy_6jfx70jq
mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit-M3 Ultra (60-core GPU) + 96GB unifiedmlx@0.31.3chat-short112.5308.7356.3r_o6wbmfegmde
Codestral-22B-v0.1-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode100.3r_4q040m4scic
Qwen2.5-32B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short71.9115.2r_twfs86tf_xf
Qwen2.5-Coder-32B-Instruct-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppconcurrent-decode71.0r_v983y0y3r2u
Qwen3-32B-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short69.5327.1r_phvxm9dcak0
gemma-2-9b-it-RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GBllama.cppchat-short69.5324.6r_1_xl4zb5-xj
qwen2.5-coderQ4_K_MRTX 3090 (24GB) + AMD EPYC 7702P 64-Core Processor (64c) + 252GBollama@0.31.1chat-short69.2402.0325.8r_6ahy-dq0f_0
qwen3.6Q4_K_MRTX 4090 (48GB) + AMD EPYC 7763 64-Core Processor (128c) + 1008GBollama@0.31.1agent-trace44.21910.51807.9r_h_659oy695r

How to read this table

Each row is the highest decode tokens-per-second measured for a unique (model, hardware) pair. When more than one workload was run, we pick the workload that produced the headline number. Lower-tps duplicates from the same machine are not shown; click through to the run page for the full per-workload breakdown.

decode tok/s is wall-clock streaming throughput (memory-bandwidth-bound). prefill tok/s is prompt ingestion throughput (compute-bound). TTFT is time-to-first-token in milliseconds. See /glossary for definitions and /methodology for the workload spec.