Skip to content
llm-speed

Phi-4 vs Gemma-2-9b-it

Signed, community-submitted decode tok/s, prefill, and TTFT for Phi-4 vs Gemma-2-9b-it — every number links to the run it came from.

Verdict

On the RTX 5090 (32GB), Gemma-2-9b-it decodes at 153 tok/s versus 141 tok/s for Phi-4, 1.1× faster. Across 3 hardware configs measured on both, Gemma-2-9b-it is faster on 3 of 3. Every cell in the table below links to the submitted run it came from.

model
Phi-4
Microsoft · 14B
View Phi-4 page →
model
Gemma-2-9b-it
Google · 9B
View Gemma-2-9b-it page →

Hardware with data on both models

HardwarePhi-4 decodeGemma-2-9b-it decodeΔSource runs
RTX 5090 (32GB)140.9tok/s152.7tok/s-11.8r_e-k4aea8ipr · r__b_bzmmab_8
M3 Ultra (60-core GPU)74.37tok/s89.45tok/s-15.1r_sqzp-0rdez- · r_iz137eqvuzy
M3 Pro (18-core GPU)15.61tok/s21.97tok/s-6.4r_-w8hnn61va_ · r_q7t7b7dcuz5

See also: Phi-4 benchmarks · Gemma-2-9b-it benchmarks · All hardware · All models · Methodology