Skip to content
llm-speed

llama3.2 on RTX 5880 (48GB) + AMD EPYC 9T24 96-Core Processor (16c) + 62GB

Workload results

WorkloadBackendModeldecode tok/sprefill tok/sTTFTp50p95
chat-shortollama@0.34.2llama3.2Q8_0466.9tok/s44.32tok/s2,821ms2.1ms2.3ms
chat-longollama@0.34.2llama3.2Q8_0441.9tok/s39,673.0tok/s79.4ms2.2ms2.4ms
concurrent-decodeollama@0.34.2llama3.2Q8_0464.2tok/s——2.1ms2.2ms
agent-traceollama@0.34.2llama3.2Q8_0452.3tok/s91,077.2tok/s21.8ms2.2ms2.4ms

Reproduce on your machine

Start with this model and workload selection. This command does not capture the original artifact, backend version or server settings:

Match quantization, context capacity, caching and runtime settings for a useful comparison. Duration varies by workload and hardware; the published client may differ from this run's version. How it's measured.

Embed this run

Drop the badge into a README, blog post, or signature. Each render is a backlink to the signed result.

llm-speed: 467 tok/s on RTX 5880 (48GB) (llama3.2)
[![llm-speed: 467 tok/s on RTX 5880 (48GB) (llama3.2)](https://llm-speed.com/badge/r_b948ehgi8g4.svg)](https://llm-speed.com/r/r_b948ehgi8g4)

Related benchmarks

Provenance

Run ID
r_b948ehgi8g4
Fingerprint hash
c8c55484fdad4e7a
Public key
bEJLrRXWo0ZHZvduikNM8kzBlCqoGiktMZ/PJgINFP8=
Received
2026-09-19 17:48:43