Skip to content
llm-speed

Qwen3-Coder-30B-A3B-Instruct vs Codestral-22B-v0.1

Signed, community-submitted decode tok/s, prefill, and TTFT for Qwen3-Coder-30B-A3B-Instruct vs Codestral-22B-v0.1 — every number links to the run it came from.

How to read these results

Selected submissions for RTX 5090 (32GB) report 260 tok/s on Qwen3-Coder-30B-A3B-Instruct and 100 tok/s on Codestral-22B-v0.1. The table selects each side's fastest submitted row independently, grouped by hardware label. Matching labels do not establish matching model artifacts, quantization, backend, workload or context. These peaks are reference results, not a controlled speed or answer-quality ranking. Inspect both source runs before choosing hardware.

model
Qwen3-Coder-30B-A3B-Instruct
Qwen · 30B-A3B
View Qwen3-Coder-30B-A3B-Instruct page →
model
Codestral-22B-v0.1
Mistral · 22B
View Codestral-22B-v0.1 page →

Hardware with data on both models

HardwareQwen3-Coder-30B-A3B-Instruct decodeCodestral-22B-v0.1 decodeΔSource runs
RTX 5090 (32GB)259.9tok/s
chat-short · llama.cpp · quantization unreported
100.3tok/s
concurrent-decode · llama.cpp · quantization unreported
+159.6r_c7qyvvmmsv1 · r_4q040m4scic
M3 Ultra (60-core GPU)112.2tok/s
chat-short · mlx · quantization unreported
47.49tok/s
chat-short · mlx · quantization unreported
+64.7r_fpsca03u2o_ · r_79dvtag5fd_

See also: Qwen3-Coder-30B-A3B-Instruct benchmarks · Codestral-22B-v0.1 benchmarks · All hardware · All models · Methodology