DeepSeek-V3 vs Llama-3.3-70B-Instruct
Signed, community-submitted decode tok/s, prefill, and TTFT for DeepSeek-V3 vs Llama-3.3-70B-Instruct — every number links to the run it came from.
Verdict
DeepSeek-V3 has no submitted run yet, so there is no head-to-head number to report. We do have 1 hardware config measured for Llama-3.3-70B-Instruct (peak 17 tok/s decode on M3 Ultra (60-core GPU)) — the table below is that baseline. Run the llm-speed suite on DeepSeek-V3 and this page fills in the second column automatically.
Llama-3.3-70B-Instruct benchmarks — no shared hardware with DeepSeek-V3 yet
| Hardware | decode tok/s | Workload | Run |
|---|---|---|---|
| M3 Ultra (60-core GPU) | 16.78tok/s | chat-short | r_sx3a4y9n-m4 |
See also: DeepSeek-V3 benchmarks · Llama-3.3-70B-Instruct benchmarks · All hardware · All models · Methodology