Skip to content
llm-speed

Llama-3.3-70B-Instruct vs Llama-3.1-405B-Instruct

Signed, community-submitted decode tok/s, prefill, and TTFT for Llama-3.3-70B-Instruct vs Llama-3.1-405B-Instruct — every number links to the run it came from.

Verdict

Llama-3.1-405B-Instruct has no submitted run yet, so there is no head-to-head number to report. We do have 1 hardware config measured for Llama-3.3-70B-Instruct (peak 17 tok/s decode on M3 Ultra (60-core GPU)) — the table below is that baseline. Run the llm-speed suite on Llama-3.1-405B-Instruct and this page fills in the second column automatically.

model
Llama-3.3-70B-Instruct
Meta · 70B
View Llama-3.3-70B-Instruct page →
model
Llama-3.1-405B-Instruct
Meta · 405B
View Llama-3.1-405B-Instruct page →

Llama-3.3-70B-Instruct benchmarks — no shared hardware with Llama-3.1-405B-Instruct yet

Hardwaredecode tok/sWorkloadRun
M3 Ultra (60-core GPU)16.78tok/schat-shortr_sx3a4y9n-m4

See also: Llama-3.3-70B-Instruct benchmarks · Llama-3.1-405B-Instruct benchmarks · All hardware · All models · Methodology