Skip to content
llm-speed

Open data

The llm-speed dataset

Download focused studies with repeated measurements and exact configuration notes, or use the dated bulk snapshot below. The bulk files contain 220 run summaries across 52 model labels and 10 reported hardware classes. Each summary links to its source run. All downloads are available under CC BY 4.0.

Bulk snapshot exported 2026-07-08 · suite suite-v1 · see the methodology for how each number is measured and signed.

Focused benchmark studies

Qwen3.8-27B and Gemma 4 12B RTX 5090 chat benchmarks — September 5, 2026

Twelve workload measurements from six accepted runs on one RTX 5090: three repetitions each of Qwen3.8-27B Q4_K_M and Gemma 4 12B QAT Q4_0, covering short and longer chat prompts. Includes decode speed, first-token latency, actual token lengths, pinned artifacts and configuration limits; no coding-quality evaluation.

Published September 12, 2026 · measured September 5 · CC BY 4.0. This study is separate from the July bulk snapshot below.

July 8 bulk snapshot

The July snapshot is also published as a HuggingFace dataset (with a live Dataset Viewer) and archived on Zenodo with a citable DOI, for notebooks, training pipelines, and papers.

Prefer more recent submissions? The JSON API returns a bounded list of run summaries (50 by default); it is not a full-corpus download. Per-run detail is at api.llm-speed.com/v1/results/{id}. Every run also has a citable page at llm-speed.com/r/{id}.

Reuse it (attribution required)

The data is licensed Creative Commons Attribution 4.0. Use it anywhere, including commercially, as long as you credit llm-speed with a link back. Copy one of these:

Data from llm-speed (https://llm-speed.com/data), CC BY 4.0.
<a href="https://llm-speed.com/data">Benchmark data by llm-speed</a>, CC BY 4.0.

Building a tool on this data (a VRAM calculator, a price-per-token comparison, a hardware guide)? That is exactly what the license is for. A link back is all we ask.

Citing it in a post or paper? Use:

llm-speed. llm-speed: signed LLM inference-speed benchmarks. July 8, 2026 snapshot. Zenodo (2026). https://doi.org/10.5281/zenodo.21254813. CC BY 4.0.

How to interpret the data

These are submitted measurements, with varying artifacts, workloads and runtime settings. A signature preserves submitted provenance; it does not independently verify hardware, answer quality or exact reproducibility. The bulk summaries omit signatures and per-workload detail. Open a source permalink before comparing rates, and preserve missing fields instead of inferring a configuration. The focused study above documents its exact cohort and limitations separately.

What is in each row

  • id and permalink:the source run and its citable page
  • top_model_name, top_backend, top_workload:what was run
  • accelerator_summary:the GPU or Apple Silicon plus host
  • top_decode_tps:the headline decode tokens per second
  • received_at, suite_version:when it landed and under which suite

The numbers are measured, not modeled. If a run looks wrong, open its permalink and check the raw workload results. Browse the same data as a leaderboard on the cheatsheet.