wrapped·Sep 19, 2026
/r/r_b948ehgi8g4your run
llama3.2
on RTX 5880
467
tok/sdecode
RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB
rank in tier
1/1RTX 5880 runs
best workload
chat-short
where the rig flew
slowest workload
chat-long442 tok/s
where the rig struggled
backend
ollamaollama is llama.cpp under the hood; pure llama.cpp tends to nudge a few % faster.
faster than
- Qwen3-0.6B-4bit on M3 Ultra429 tok/s
- stable-code-instruct-3b on RTX 5090356 tok/s
- gpt-oss-20b on RTX 5090318 tok/s