wrapped·Sep 13, 2026
/r/r_b1rn5gi6b3byour run
llama3.2
23.9
tok/sdecode
AMD64 Family 23 Model 104 Stepping 1, AuthenticAMD (6c) + 15GB
rank in tier
1/1AMD64 Family 23 Model 104 Stepping 1, AuthenticAMD runs
best workload
chat-short
where the rig flew
slowest workload
chat-long14.7 tok/s
where the rig struggled
backend
ollamaollama is llama.cpp under the hood; pure llama.cpp tends to nudge a few % faster.
faster than
- gemma-2-9b-it-4bit on M3 Pro22.0 tok/s
- stable-code-instruct-3b-4bit on M3 Pro19.4 tok/s
- llama-3.3-70b-Instruct-4bit on M3 Ultra16.8 tok/s