Skip to content
llm-speed
Leaderboard/models/deepseek-r1

r1

16 workload results across 1 hardware configuration.

Fastest local config

133.8 decode tok/s

on RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB via ollama (Q4_K_M). see full run

Local runs (16 runs)

Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.

RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GBRTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB

WorkloadBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-shortollama@0.31.1Q4_K_Mno datano datano datar_zrjj-pj93q8
chat-longollama@0.31.1Q4_K_M81.09tok/s240.4tok/s13,070msr_zrjj-pj93q8
concurrent-decodeollama@0.31.1Q4_K_Mno datano datano datar_zrjj-pj93q8
agent-traceollama@0.31.1Q4_K_Mno datano datano datar_zrjj-pj93q8
chat-shortollama@0.31.1Q4_K_Mno datano datano datar_fg77v2hhohb
chat-longollama@0.31.1Q4_K_M131.9tok/s686.3tok/s4,578msr_fg77v2hhohb
concurrent-decodeollama@0.31.1Q4_K_Mno datano datano datar_fg77v2hhohb
agent-traceollama@0.31.1Q4_K_M133.8tok/s8,852.4tok/s366msr_fg77v2hhohb
chat-shortollama@0.31.1Q4_K_Mno datano datano datar_l2ck67ls40i
chat-longollama@0.31.1Q4_K_M3.76tok/s14.28tok/s219,986msr_l2ck67ls40i
concurrent-decodeollama@0.31.1Q4_K_Mno datano datano datar_l2ck67ls40i
agent-traceollama@0.31.1Q4_K_M3.79tok/s3,515.1tok/s971msr_l2ck67ls40i
chat-shortollama@0.31.1Q4_K_Mno datano datano datar_2r1p3w9ps7c
chat-longollama@0.31.1Q4_K_M131.0tok/s599.9tok/s5,238msr_2r1p3w9ps7c
concurrent-decodeollama@0.31.1Q4_K_Mno datano datano datar_2r1p3w9ps7c
agent-traceollama@0.31.1Q4_K_M132.4tok/s9,612.8tok/s337msr_2r1p3w9ps7c

Community folklore

102 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.

  • communityconfidence 70%

    2.13tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.00tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.13tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.00tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.13tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.00tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.13tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 70%

    2.00tok/s DeepSeek R1 on RTX 5090 via llama.cpp

    our signed data: RTX 5090 · DeepSeek R1

    DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t

    source: Reddit · u/VoidAlchemy · 2025-01-30

  • communityconfidence 60%

    18.43tok/s DeepSeek R1 on M3 Ultra via mlx

    our signed data: M3 Ultra · DeepSeek R1

    R1 671B Q4 - M3 Ultra 512GB with MLX🔥 Yes it works! First test, and I'm blown away! Prompt: "Create an amazing animation using p5js" * **18.43 tokens/sec** * Generates a p5js zero-shot, tested at video's end * Video in real-time, no acceleration! https://reddit.com/link/1j9vj…

    source: Reddit · u/ifioravanti · 2025-03-12

  • communityconfidence 60%

    7.00tok/s deepseek-r1 on M3 Max via ollama

    our signed data: M3 Max · deepseek-r1

    oubles the electricity cost compared to RTX 6000 ADA (19tps) or RTX A6000 (12tps). \- \`M3 Max 40GPU\` has high memory but only delivers 3-7 tps for \`deepseek-r1:70b\`. It is also loud, and the GPU temperature is high (> 90 C). https://preview.redd.it/8r7cwajfn9ee1.png?width=1…

    source: Reddit · u/Joehua87 · 2025-01-21

See all 102 claims for r1

r1 on hardware