How fast is GLM-4.7-Flash on an RTX 4090?
Published 2026-07-08 · updated 2026-09-16
The cited RTX 4090 result reports 129.9 tok/s for the final timed turn of an agent trace. It does not establish a general chat speed or a coding-quality score. The July 8, 2026 submission used GLM-4.7-Flash Q4_K_M through Ollama 0.31.1 on a host reporting a 24GB RTX 4090, an AMD EPYC 75F3 and 504GB of system RAM.
If you are deciding whether to run it on your own 4090, check three things: the workload behind the speed number, the model's actual GPU placement, and the context your application needs. This guide separates the recorded result from what still needs testing.
What the RTX 4090 source actually records
Open the original GLM submission. It contains four workload rows from suite-v1 / CLI 0.0.3. The short-chat, long-chat and concurrent-decode rows have no primary decode or first-token measurement. Only the agent-trace row has those fields populated.
| Recorded workload | Input tokens | Output tokens | Decode tok/s | First token |
|---|---|---|---|---|
| Short chat | 105 | 256 | Missing | Missing |
| Longer chat | 3,132 | 1,024 | Missing | Missing |
| Final timed agent turn | 3,145 | 120 | 129.9 | 1,262.5 ms |
The trace contains seven assistant turns; six lack per-turn timing. Its top-level output count is 840, while the seven per-turn output counts sum to 552. The table uses the final turn's own 120-token count. We do not turn that inconsistent total into a whole-trace throughput score. Backend timing fields also exist for the chat workloads, but they are a different measurement source and are not substituted for the missing primary values.
The 1,262.5 ms figure is the recorded first-token wait, not an isolated prompt-processing time. The source includes model-loading time. There are no repeated trials or coding correctness scores here, and roughly 3,145 input tokens do not test a 32K or 128K prompt. See decode, prefill and first-token latency for the distinction.
Does GLM-4.7-Flash fit a 24GB RTX 4090?
Ollama currently lists the explicit Q4_K_M tag at about 19GB. That is the download size, not the full GPU allocation. Weights, context cache, runtime buffers and other applications all consume memory. The historical run does not record peak VRAM or the CPU/GPU placement split, so it cannot prove full GPU residency or spare memory at your context size.
Z.ai describes GLM-4.7-Flash as a 30B-A3B mixture-of-experts model. The active-parameter label is not a 3B-sized download. Likewise, the model listing's 198K context limit is not evidence that this 24GB card can allocate that much context.Ollama documents that larger context increases memory use. Start with the context your task actually needs, then inspect the loaded model with ollama ps.
Is it slower than Qwen3-Coder on a 4090?
The previously cited 179.9 tok/s Qwen3-Coder result comes from concurrent-decode, whereas GLM's 129.9 comes from an agent-trace turn. The Qwen host reports an EPYC 7352 and 252GB RAM, and its agent-trace workload failed. Both report an RTX 4090, Q4_K_M and Ollama 0.31.1, but this is not a matched speed comparison.
These two values cannot establish a percentage speed advantage or which model solves your coding tasks better. Compare the same prompts, artifact precision, context, cache state and concurrency, and score whether the answers work. For a newer measured configuration on different hardware, see our Qwen3.8 versus Gemma 4 RTX 5090 study; those results do not transfer directly to the 4090.
Check your configuration before buying hardware
- Record the exact model tag and digest, Ollama version, quantization, context setting and whether thinking is enabled. A current tag may differ from the July artifact.
- While the model is loaded, run
ollama ps. CheckPROCESSORandCONTEXT; Ollama explains the CPU/GPU split. Reduce unnecessary context or competing model residency before assuming you need a new GPU. - Measure your real prompt length, first-token wait, generation time and answer correctness. Record cold/loading behavior separately from subsequent requests.
With llm-speed installed, Ollama running and the explicit Q4_K_M model already downloaded, this selects the agent-trace workload and saves a local result for inspection:
llm-speed bench --backend ollama \
--model glm-4.7-flash:q4_K_M --workload agent-trace \
--strict-anon --no-upload --json glm-agent-trace.jsonThis is a starting configuration, not a claim of exact reproduction of the July run. The historical record lacks a pinned model artifact and complete server settings; a newer client may also measure differently. Inspect the result before deciding to contribute it. Thecontribution guide covers installation and submission. If your run is unexpectedly slow, follow the GPU offload and context checklist.