How fast is Qwen3.6-27B on an RTX 4090?
Published 2026-07-03 · updated 2026-09-16
Our historical Qwen3.6 records report 44.18 tok/s on a host labeled RTX 4090 (48GB) and 73.96 tok/s on an RTX 5090 (32GB). They use different workloads and runtimes. The 4090 record does not establish performance or memory fit on a standard single 24 GB card, and dividing these speeds would not measure the benefit of a GPU upgrade.
What the three linked runs actually measured
| Model label | Reported GPU memory | Runtime / workload | Decode tok/s |
|---|---|---|---|
| Qwen3.6, 27.8B, Q4_K_M | RTX 4090 · 48 GB | Ollama 0.31.1 · agent-trace | 44.18 |
| Qwen3.6-27B-Q4_K_M.gguf | RTX 5090 · 32 GB | llama.cpp · chat-short | 73.96 |
| Qwen3.6-35B-A3B-Q4_K_M.gguf | RTX 5090 · 32 GB | llama.cpp · chat-short | 223.97 |
The 4090 submission describes an EPYC host with 1,008 GB of system memory. Its 48 GB accelerator summary does not tell us the physical card arrangement or establish how much of the model was GPU-resident. The displayed speed is the populated agent-trace result; its recorded first-token wait is about 1.81 seconds for the last measured assistant turn, not the whole trace. The short-chat, long-chat and concurrent rows have no top-level decode value.
Both 5090 submissions describe the same Ryzen host, use the short-chat workload and record 256 output tokens. Their first-token waits are about 190 ms for the dense 27B and 167 ms for the 35B-A3B. However, the runtime version is missing, prompt/context counts are stored as zero, and the Q4_K_M designation appears in the filename rather than a populated quantization field. Zero here does not establish that an empty prompt was measured. These records cannot support a controlled hardware speedup or coding-quality ranking.
Will Qwen3.6-27B fit a 24 GB RTX 4090?
The cited 48 GB record cannot settle that question. Check the exact downloaded weights, GPU allocation, runtime buffers and context-cache memory at your intended prompt length. Keep the model loaded while testing the longest request and any simultaneous users you expect. A Q4 label alone is not a guarantee of fit or spare context capacity.
The memory-fit calculator can screen configurations, but its estimates are not a measured Qwen3.6 fit test. For a purchase decision, use the 4090-to-5090 upgrade checklist and compare a matched workload before buying.
Dense 27B or 35B-A3B for a coding agent?
Qwen publishes separate 27B and 35B-A3B model cards. The latter is a mixture-of-experts model: activating fewer parameters per token does not mean only those parameters must fit in memory. Publisher evaluations, local decode rates and success on your own repository tasks answer different questions.
The 35B-A3B record has a higher observed streaming rate, but these runs do not show which model solves more coding tasks. Compare both on a small set of your real edits, including tool calls and tests. Record correct completions, first-token delay and total time, with the same prompt and output budget.
Reproduce a useful comparison
Follow the benchmark setup guide to select the intended model and workload. Record the exact artifact, quantization, runtime version, offload settings, configured context and actual token counts; repeat warm measurements before summarizing a range. Installing the CLI alone does not select either Qwen3.6 variant or reproduce these historical settings.
For newer evidence, our Qwen3.8-27B and Gemma 4 RTX 5090 study records pinned files and three repetitions per model. It is a separate experiment, with no matched 4090 run or coding-quality comparison.
Source review: September 16, 2026. These are historical submitted measurements, received July 1–2; the review date is not a new benchmark date. Linked run records remain the source for the reported values and limitations.