Open data
The llm-speed dataset
Download focused studies with repeated measurements and exact configuration notes, or use the dated bulk snapshot below. The bulk files contain 220 run summaries across 52 model labels and 10 reported hardware classes. Each summary links to its source run. All downloads are available under CC BY 4.0.
Bulk snapshot exported 2026-07-08 · suite suite-v1 · see the methodology for how each number is measured and signed.
Focused benchmark studies
Qwen3.8-27B and Gemma 4 12B RTX 5090 chat benchmarks — September 5, 2026
Twelve workload measurements from six accepted runs on one RTX 5090: three repetitions each of Qwen3.8-27B Q4_K_M and Gemma 4 12B QAT Q4_0, covering short and longer chat prompts. Includes decode speed, first-token latency, actual token lengths, pinned artifacts and configuration limits; no coding-quality evaluation.
Published September 12, 2026 · measured September 5 · CC BY 4.0. This study is separate from the July bulk snapshot below.
July 8 bulk snapshot
The July snapshot is also published as a HuggingFace dataset (with a live Dataset Viewer) and archived on Zenodo with a citable DOI, for notebooks, training pipelines, and papers.
Prefer more recent submissions? The JSON API returns a bounded list of run summaries (50 by default); it is not a full-corpus download. Per-run detail is at api.llm-speed.com/v1/results/{id}. Every run also has a citable page at llm-speed.com/r/{id}.
Reuse it (attribution required)
The data is licensed Creative Commons Attribution 4.0. Use it anywhere, including commercially, as long as you credit llm-speed with a link back. Copy one of these:
Data from llm-speed (https://llm-speed.com/data), CC BY 4.0.<a href="https://llm-speed.com/data">Benchmark data by llm-speed</a>, CC BY 4.0.Building a tool on this data (a VRAM calculator, a price-per-token comparison, a hardware guide)? That is exactly what the license is for. A link back is all we ask.
Citing it in a post or paper? Use:
llm-speed. llm-speed: signed LLM inference-speed benchmarks. July 8, 2026 snapshot. Zenodo (2026). https://doi.org/10.5281/zenodo.21254813. CC BY 4.0.How to interpret the data
These are submitted measurements, with varying artifacts, workloads and runtime settings. A signature preserves submitted provenance; it does not independently verify hardware, answer quality or exact reproducibility. The bulk summaries omit signatures and per-workload detail. Open a source permalink before comparing rates, and preserve missing fields instead of inferring a configuration. The focused study above documents its exact cohort and limitations separately.
What is in each row
idandpermalink:the source run and its citable pagetop_model_name,top_backend,top_workload:what was runaccelerator_summary:the GPU or Apple Silicon plus hosttop_decode_tps:the headline decode tokens per secondreceived_at,suite_version:when it landed and under which suite
The numbers are measured, not modeled. If a run looks wrong, open its permalink and check the raw workload results. Browse the same data as a leaderboard on the cheatsheet.