Skip to content
llm-speed

Rtx 5880 LLM benchmark

The highest recorded decode rate in these Rtx 5880 submissions is from llama3.2 (1.2B) at 466.9 decode tok/s via ollama (source run). This page contains 4 workload rows from 1 submitted run across 1 model labels. Model sizes, workloads and reported hardware configurations differ; the highest rate is not a quality ranking or a speed guarantee.

Highest recorded decode rate in these submissions

466.9 decode tok/s

llama3.2 (1.2B) via ollama (Q8_0). Workload: chat-short. Reported hardware: RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB. see full run

llama3.2

WorkloadModel sizeReported hardwareBackendQuantdecode tok/sprefill tok/sTTFTRun
chat-short1.2B
RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB
ollama@0.34.2Q8_0466.9tok/s44.32tok/s2,821msr_b948ehgi8g4
chat-long1.2B
RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB
ollama@0.34.2Q8_0441.9tok/s39,673.0tok/s79.4msr_b948ehgi8g4
concurrent-decode1.2B
RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB
ollama@0.34.2Q8_0464.2tok/sno datano datar_b948ehgi8g4
agent-trace1.2B
RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB
ollama@0.34.2Q8_0452.3tok/s91,077.2tok/s21.8msr_b948ehgi8g4

Models measured on Rtx 5880

Common questions about Rtx 5880

Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.

Read the Rtx 5880 FAQ →