Skip to content
llm-speed

llama3.2 on AMD64 Family 23 Model 104 Stepping 1, AuthenticAMD (6c) + 15GB

Workload results

WorkloadBackendModeldecode tok/sprefill tok/sTTFTp50p95
chat-shortollama@0.34.0llama3.2Q8_023.88tok/s39.41tok/s3,172ms40.8ms49.3ms
chat-longollama@0.34.0llama3.2Q8_014.69tok/s93.91tok/s33,563ms65.7ms92.5ms
concurrent-decodeollama@0.34.0llama3.2Q8_019.96tok/s——48.9ms61.0ms
agent-traceollama@0.34.0llama3.2Q8_016.80tok/s203.2tok/s7,977ms59.1ms80.2ms

Reproduce on your machine

Start with this model and workload selection. This command does not capture the original artifact, backend version or server settings:

Match quantization, context capacity, caching and runtime settings for a useful comparison. Duration varies by workload and hardware; the published client may differ from this run's version. How it's measured.

Embed this run

Drop the badge into a README, blog post, or signature. Each render is a backlink to the signed result.

llm-speed: 23.9 tok/s on AMD64 Family 23 Model 104… (llama3.2)
[![llm-speed: 23.9 tok/s on AMD64 Family 23 Model 104… (llama3.2)](https://llm-speed.com/badge/r_b1rn5gi6b3b.svg)](https://llm-speed.com/r/r_b1rn5gi6b3b)

Related benchmarks

Provenance

Run ID
r_b1rn5gi6b3b
Fingerprint hash
ad5a7e35f4d4cd2d
Public key
RLSp8LxZhTgXzXA05NEeNvk040OpZay0waz2cCjUmB8=
Received
2026-09-13 18:33:05