stable-code-instruct-3b-4bit
7 workload results across 3 hardware configurations.
Fastest local config
356.1 decode tok/s
on RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB via llama.cpp. see full run
Local runs (7 workload results)
Runs from contributors' own machines via MLX, llama.cpp, vLLM, exllamav2, or ollama. Signed on the submitter's hardware.
RTX 5090 (32GB) + AMD Ryzen 7 9850X3D 8-Core Processor (8c) + 30GB
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | llama.cpp | - | 341.9tok/s | no data | 175ms | r_q9f15lz6831 |
| chat-long | llama.cpp | - | 289.1tok/s | no data | 232ms | r_q9f15lz6831 |
| concurrent-decode | llama.cpp | - | 356.1tok/s | no data | no data | r_q9f15lz6831 |
| agent-trace | llama.cpp | - | 292.8tok/s | 44,387.5tok/s | 42.7ms | r_q9f15lz6831 |
| chat-short | llama.cpp | - | 331.0tok/s | no data | 36.7ms | r_8l57cim1i10 |
M3 Ultra (60-core GPU) + 96GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 192.5tok/s | 560.7tok/s | 226ms | r_y2_5y8oo97d |
M3 Pro (18-core GPU) + 36GB unified
| Workload | Backend | Quant | decode tok/s | prefill tok/s | TTFT | Run |
|---|---|---|---|---|---|---|
| chat-short | mlx@0.31.3 | - | 19.37tok/s | 131.3tok/s | 967ms | r_pqjsvd-cub4 |