llm-speed (2026-09-06). Gemma 4 12B IT QAT Q4_0 on RTX 5090: long-context measurements. https://llm-speed.com/m/gemma-4-12b-it-qat#long-context Figure: https://llm-speed.com/data/gemma-4-5090-context-2026-09-06-figure.png Data: https://llm-speed.com/data/gemma-4-5090-context-2026-09-06.json License: CC BY 4.0 — https://creativecommons.org/licenses/by/4.0/ Nine measurements: three sequential repetitions at each actual prompt length (23,872 / 47,652 / 95,212 tokens), not estimated 32K/64K/128K targets. Warm model, one active request, prompt caching disabled, thinking off. llama.cpp CUDA revision 9725a313be0528214c4a02fed906ddaf7b3f712e. Dots are medians; bars are observed minima and maxima, not confidence intervals. No full 128K-input, maximum-context, multi-user throughput or answer-quality claim. Source runs: https://llm-speed.com/r/r_8q-uf-rq-gp https://llm-speed.com/r/r_1srqmpjnkgd https://llm-speed.com/r/r_-lh-vczf7ik When adapting the figure, retain attribution and identify your changes.