Blog
Local LLM speed, measured
Benchmarks, setup guides and exploratory tests, with source links and limitations. Articles distinguish signed submissions, publisher reports and our own exploratory findings.
- Qwen3.8-27B on M5 Max 48GB: MLX speed and context latency2026-09-16
Original Qwen3.8-27B 4-bit MLX study on a 48GB M5 Max: 12 samples, first-token waits, prefix-cache costs, memory observations and downloadable reproduction files.
Read - Gemma 4 E2B JSON output: valid JSON can still route incorrectly2026-09-15
RTX 5090 test of Gemma 4 E2B QAT Q4_0 in llama.cpp: schema constraints fixed JSON parsing in nine small cases, but one routing decision remained wrong. Raw results and settings.
Read - Why is my local LLM slow? Check GPU offload, context and latency2026-09-12
Diagnose slow local LLMs: GPU placement, SGLang context limits, Ollama memory, multiple models and queued requests. Check the setup before buying hardware.
Read - Qwen3.8-27B vs Gemma 4 12B: RTX 5090 benchmarks2026-09-12
Measured RTX 5090 comparison: Qwen3.8-27B Q4_K_M vs Gemma 4 12B QAT Q4_0. Three repeats, first-token latency, GPU memory, exact files and downloadable results.
Read - What hardware do you need to run DeepSeek-V4-Flash?2026-07-08
DeepSeek V4 local hardware: documented RTX 5090 plus system-RAM serving, runtime limits, and why the new V4.1-Flash checkpoint does not fit one 32GB GPU.
Read - How fast is GLM-4.7-Flash on an RTX 4090?2026-07-08
GLM-4.7-Flash on RTX 4090: inspect the 129.9 tok/s agent-trace result, missing chat timings, 24GB memory limits and a local benchmark setup.
Read - How fast is Gemma 3 on an RTX 4090? (4B, 12B, 27B)2026-07-03
Gemma 3 4B, 12B and 27B RTX 4090 source runs: short and longer prompt speeds, first-token waits, memory limits and how to choose a configuration.
Read - Best models for a 512GB Mac Studio: speed and memory2026-07-03
512GB Mac Studio guide: M5 Ultra availability, pinned Qwen3.8 and GLM-5.3 downloads, memory checks, and clearly labeled M3 Ultra benchmark evidence.
Read - How fast is Qwen3.6-27B on an RTX 4090?2026-07-03
Qwen3.6 benchmark records report 44 tok/s on a host labeled RTX 4090 48GB and 74 on a 5090. Inspect workload differences, memory-fit limits and coding tradeoffs.
Read - Local LLM inference speed in 2026: what we've measured2026-07-02
Signed, reproducible decode tok/s across consumer GPUs and Apple Silicon: the fastest configs, why small-MoE coders win, and what counts as fast enough.
Read - The fastest local coding models in 2026 (measured on an RTX 5090)2026-07-02
Signed, reproducible decode tok/s for the top local coding models on a single RTX 5090, and why small-MoE coders decode ~4x faster than a dense 32B.
Read - The RTX 3090 in 2026: still the value pick for local LLMs?2026-07-02
Signed benchmarks on an RTX 3090 (24 GB): the decode tok/s a used ~$700 card really delivers for local LLMs in 2026. Measured, not folklore.
Read - Decode vs prefill tok/s: what LLM speed numbers actually mean2026-07-02
Prefill, decode and TTFT explained with measured Gemma 4 and Qwen results: fast streaming, long waits, and the preparation cost outside cached request timing.
Read - RTX 3090 vs RTX 4090 for local LLMs: speed and 24GB limits2026-07-02
Compare RTX 3090 and 4090 LLM source runs, first-token delays, 24GB memory limits and upgrade criteria. Reported 48GB hosts prevent a standard-card speed claim.
Read - What's the fastest GPU for running Llama 3.1 8B locally?2026-07-02
Measured decode tok/s for Llama 3.1 8B on the RTX 5090, 4090, 3090, and Apple Silicon: which runs it fastest, and which is the best value.
Read - The best GPU for a local coding agent in 20262026-07-02
Choose hardware for a local coding agent using task quality, memory and latency. Historical 3090/4090 source rows, 5090 studies and Mac configuration caveats.
Read - RTX 5090 vs M3 Ultra for local LLMs: speed vs memory2026-07-02
Compare recorded RTX 5090 and 96GB M3 Ultra short-chat speed, first-token latency and memory limits, with source runs and clear configuration caveats.
Read - How fast is Qwen3-Coder locally? RTX 5090, 4090, and Apple Silicon2026-07-02
Signed decode benchmarks for Qwen3-Coder-30B-A3B: 260 tok/s on an RTX 5090, 180 on a 4090, and ~112 on Apple Silicon. Why a small-MoE coder is fast everywhere.
Read - Codestral vs Qwen3-Coder: which local coding model is faster?2026-07-02
Signed benchmarks: Qwen3-Coder-30B-A3B decodes about 2.5x faster than Codestral-22B on the same hardware, because it is a small mixture-of-experts. Which to run.
Read - DeepSeek-R1 on an RTX 4090: which size actually fits?2026-07-02
DeepSeek-R1 8B, 14B and 32B RTX 4090 source runs, memory-fit limits and a GPU-placement checklist. A slow 32B result does not prove you need a bigger card.
Read