Rtx 5880 LLM benchmark
The highest recorded decode rate in these Rtx 5880 submissions is from llama3.2 (1.2B) at 466.9 decode tok/s via ollama (source run). This page contains 4 workload rows from 1 submitted run across 1 model labels. Model sizes, workloads and reported hardware configurations differ; the highest rate is not a quality ranking or a speed guarantee.
Highest recorded decode rate in these submissions
466.9 decode tok/s
llama3.2 (1.2B) via ollama (Q8_0). Workload: chat-short. Reported hardware: RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB. see full run
llama3.2
| Workload | Model size | Reported hardware | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|---|---|
| chat-short | 1.2B | RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB | ollama@0.34.2 | Q8_0 | 466.9tok/s | 44.32tok/s | 2,821ms | r_b948ehgi8g4 |
| chat-long | 1.2B | RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB | ollama@0.34.2 | Q8_0 | 441.9tok/s | 39,673.0tok/s | 79.4ms | r_b948ehgi8g4 |
| concurrent-decode | 1.2B | RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB | ollama@0.34.2 | Q8_0 | 464.2tok/s | no data | no data | r_b948ehgi8g4 |
| agent-trace | 1.2B | RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB | ollama@0.34.2 | Q8_0 | 452.3tok/s | 91,077.2tok/s | 21.8ms | r_b948ehgi8g4 |
Models measured on Rtx 5880
Common questions about Rtx 5880
Direct Q&A drawn from the runs above: fastest LLM, supported model classes, backend rankings, quantization guidance.