r1
16 workload results across 1 hardware configuration.
Fastest local config
133.8 decode tok/s
on RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB via ollama (Q4_K_M). see full run
Local runs (16 runs)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
RTX 4090 (24GB) + AMD EPYC 75F3 32-Core Processor (64c) + 504GB
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_zrjj-pj93q8 |
| chat-long | ollama@0.31.1 | Q4_K_M | 81.09tok/s | 240.4tok/s | 13,070ms | r_zrjj-pj93q8 |
| concurrent-decode | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_zrjj-pj93q8 |
| agent-trace | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_zrjj-pj93q8 |
| chat-short | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_fg77v2hhohb |
| chat-long | ollama@0.31.1 | Q4_K_M | 131.9tok/s | 686.3tok/s | 4,578ms | r_fg77v2hhohb |
| concurrent-decode | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_fg77v2hhohb |
| agent-trace | ollama@0.31.1 | Q4_K_M | 133.8tok/s | 8,852.4tok/s | 366ms | r_fg77v2hhohb |
| chat-short | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_l2ck67ls40i |
| chat-long | ollama@0.31.1 | Q4_K_M | 3.76tok/s | 14.28tok/s | 219,986ms | r_l2ck67ls40i |
| concurrent-decode | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_l2ck67ls40i |
| agent-trace | ollama@0.31.1 | Q4_K_M | 3.79tok/s | 3,515.1tok/s | 971ms | r_l2ck67ls40i |
| chat-short | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_2r1p3w9ps7c |
| chat-long | ollama@0.31.1 | Q4_K_M | 131.0tok/s | 599.9tok/s | 5,238ms | r_2r1p3w9ps7c |
| concurrent-decode | ollama@0.31.1 | Q4_K_M | no data | no data | no data | r_2r1p3w9ps7c |
| agent-trace | ollama@0.31.1 | Q4_K_M | 132.4tok/s | 9,612.8tok/s | 337ms | r_2r1p3w9ps7c |
Community folklore
102 unverified claims extracted from Reddit/HN comments. Lower trust than signed runs above; every row links to the source.
- communityconfidence 70%
2.13tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…”
- communityconfidence 70%
2.00tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t”
- communityconfidence 70%
2.13tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…”
- communityconfidence 70%
2.00tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t”
- communityconfidence 70%
2.13tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…”
- communityconfidence 70%
2.00tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t”
- communityconfidence 70%
2.13tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“ok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 tok/sec with 2k context using a dynamic quant of the full R1 671B model (not a distill) after *disabling* my 3090TI GPU on a 96GB RAM g…”
- communityconfidence 70%
2.00tok/s — DeepSeek R1 on RTX 5090 via llama.cpp
our signed data: RTX 5090 · DeepSeek R1
“DeepSeek R1 671B over 2 tok/sec *without* GPU on local gaming rig! Don't rush out and buy that 5090TI just yet (if you can even find one lol)! I just inferenced \~2.13 t”
- communityconfidence 60%
18.43tok/s — DeepSeek R1 on M3 Ultra via mlx
our signed data: M3 Ultra · DeepSeek R1
“R1 671B Q4 - M3 Ultra 512GB with MLX🔥 Yes it works! First test, and I'm blown away! Prompt: "Create an amazing animation using p5js" * **18.43 tokens/sec** * Generates a p5js zero-shot, tested at video's end * Video in real-time, no acceleration! https://reddit.com/link/1j9vj…”
- communityconfidence 60%
7.00tok/s — deepseek-r1 on M3 Max via ollama
our signed data: M3 Max · deepseek-r1
“oubles the electricity cost compared to RTX 6000 ADA (19tps) or RTX A6000 (12tps). \- \`M3 Max 40GPU\` has high memory but only delivers 3-7 tps for \`deepseek-r1:70b\`. It is also loud, and the GPU temperature is high (> 90 C). https://preview.redd.it/8r7cwajfn9ee1.png?width=1…”