Skip to content
llm-speed

DeepSeek-V3 vs Llama-3.3-70B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for DeepSeek-V3 vs Llama-3.3-70B-Instruct — every number links to the run it came from.

Verdict

DeepSeek-V3 has no submitted run yet, so there is no head-to-head number to report. We do have 1 hardware config measured for Llama-3.3-70B-Instruct (peak 17 tok/s decode on M3 Ultra (60-core GPU)) — the table below is that baseline. Run the llm-speed suite on DeepSeek-V3 and this page fills in the second column automatically.

model
DeepSeek-V3
DeepSeek · 671B-A37B
View DeepSeek-V3 page →
model
Llama-3.3-70B-Instruct
Meta · 70B
View Llama-3.3-70B-Instruct page →

Llama-3.3-70B-Instruct benchmarks — no shared hardware with DeepSeek-V3 yet

Hardwaredecode tok/sWorkloadRun
M3 Ultra (60-core GPU)16.78tok/schat-shortr_sx3a4y9n-m4

See also: DeepSeek-V3 benchmarks · Llama-3.3-70B-Instruct benchmarks · All hardware · All models · Methodology