Qwen3.8-27B on M5 Max 48GB: MLX speed and context latency
Published 2026-09-16
Qwen3.8-27B ran on our 48 GiB M5 Max, but a longer prompt meant a much longer wait before the first token. With the tested MLX 4-bit conversion, the median wait rose from 8.17 seconds at 840 input tokens to 113.21 seconds at 12,263 tokens. Once output began, median stream decode was about 13–14 tokens per second across those fresh-context conditions.
These are September 16 measurements on a shared workstation with uncontrolled background load. Timing varied substantially: the shortest input ranged from 11.48 to 27.17 decode tok/s. Treat the results as a documented observation of this setup, not the expected speed of every M5 Max or a comparison with Ollama.
All 12 measurements (CSV) · Prompts, outputs and timing events (JSON) · Evidence and reproduction bundle (ZIP)
How fast was Qwen3.8-27B on the M5 Max?
The Mac reported 40 GPU cores and 48 GiB of unified memory. The model was already loaded, thinking was disabled, and each baseline request started with a fresh prompt cache. Every response reached the 256-token output cap. Three repetitions per input are shown; parentheses are the observed minimum–maximum, not confidence intervals.
| Actual input tokens | First token (s) | Stream decode (tok/s) | Peak MLX allocation (GiB) |
|---|---|---|---|
| 840 | 8.17 (3.79–11.78) | 13.70 (11.48–27.17) | 15.67 |
| 3,156 | 27.66 (16.59–35.35) | 13.62 (11.09–25.92) | 16.96 |
| 12,263 | 113.21 (74.39–133.19) | 13.34 (10.10–18.22) | 18.48 |
Stream decode measures the 255 intervals after the first generated token. The raw files also retain the runtime's slightly different generation-rate calculation. First visible text and first generated token coincided in these samples. Peak MLX allocation includes model and MLX working allocations; it is not total system RAM or a minimum installed-memory recommendation.

Does prefix caching remove the long wait?
It moved substantial work before the request timer. We separately prepared 12,199 tokens of the same 12,263-token input, then supplied the remaining 64 tokens. The cached request's first-token wait was 0.87 (0.71–1.34) seconds, but prefix preparation took 120.80 (118.65–134.77) seconds.
Counting preparation and first-token wait together, the three cached trials took 119.51–136.11 seconds. There is no free speedup for a one-off prompt. Reusing an already prepared prefix can make a later request start sooner, but this experiment rebuilt it for every pair and did not measure an actual multi-turn agent or automatic cache reuse in an application.
Cached stream decode was 12.25 (11.05–14.48) tok/s. The fresh and cached responses were not token-identical: all three pairs first diverged at zero-based output index 188. We did not evaluate which response was better or establish the cause. A faster cached request therefore does not establish identical behavior.
Did it fit without memory pressure?
The completed study's peak MLX allocation reached 18.48 GiB, but that does not prove a smaller Mac would behave the same. Before the successful run, a first smoke attempt was stopped when system swap usage grew by about 7.16 GiB. It completed no measured requests. We preserved that failed attempt in the download rather than treating the successful retry as the whole story.
After headroom recovered, one smoke retry and the 12-sample study completed. The main study started with about 12.45 GiB of existing system swap usage and showed no additional growth in the sampled swap-used counter; that does not establish zero paging activity. The lowest sampled system-wide free-memory percentage was 23%. Background applications remained uncontrolled, and thermal-limit readings were unavailable.
For a machine you already own, test your actual prompt and inspect memory pressure before deciding whether to upgrade. See the Mac memory and hardware checklist. These results cover neither M5 Pro nor 128K input, maximum context, concurrent users, image input or coding-task success.
Exact model and measurement method
- Model: mlx-community/Qwen3.8-27B-4bit, pinned revision 3e6447f. All 14 downloaded files were verified; the three weight shards total 16,054,541,349 bytes. Full hashes are in the bundle. This is a distinct artifact from the GGUF in our RTX 5090 study.
- Runtime: MLX 0.32.1, mlx-lm 0.31.3, transformers 5.17.0, tokenizers 0.23.2, huggingface-hub 1.31.0, NumPy 2.5.2, Python 3.14.7; macOS 26.6.2, AC power.
- Text-only synthetic document summaries; greedy sampling, seed 0, batch one, thinking off, prefill chunks of 2,048 tokens. No KV quantization, rotating cache or draft model. The separate warmup is excluded.
- Three repetitions in ascending input-length order, each followed by the long prepared-prefix pair. The order was not randomized. Later fresh requests were slower; background load and timing cannot be separated into causal effects here.
- The prompt requests at least 400 words, but all outputs are truncated at 256 tokens. This measures generation under that cap, not finished summaries or model quality.
- Model loading took 2.69 seconds and is excluded from request timing. Tokenization and cached-prefix preparation are also outside the request timer. Cached peak allocation includes prefix preparation.
The supervisor sampled memory every three seconds and stopped only its own process if free memory fell below 20%, additional system swap exceeded 512 MiB, process RSS exceeded 24 GiB, or wall time exceeded 1,250 seconds. The child had a 1,200-second timeout. MLX's 22 GiB memory setting was a guideline, not a hard cap; normal runtime wired-memory behavior was unchanged. Sampling does not bound between-sample peaks.
This is an exploratory native-MLX study, separate from signed suite-v1 leaderboard submissions. The nominal fixture labels in the raw file are character-based targets; the table uses the tokenizer's actual counts of 840, 3,156 and 12,263.
Inspect, reproduce or cite the study
Download the 130 KB evidence bundle for the exact executed runner, fixture, pinned primary packages, successful and failed smoke evidence, monitor samples, per-token events and checksums. No weights, credentials, personal paths or machine serials are included. The included README documents setup and the stronger memory gate used before the successful runs. Transitive dependencies are not fully locked, so identical timing in another environment is not guaranteed.
To validate the saved measurements without downloading a model, extract a copy and run python3 summarize_study.py. To repeat inference on a suitable Apple Silicon Mac, follow the README to create the isolated runtime, verify the pinned download, pass the supervised smoke test and then run the study. Use a separate directory: the scripts overwrite their output files.
Measurements and figures are available under CC BY 4.0; the runner and fixture use Apache-2.0. Cite: llm-speed, Qwen3.8-27B on M5 Max with 48 GiB, 16 September 2026
, link this page, and keep the shared-workstation and configuration limitations with the numbers.
What should you check in a slow local agent?
Record time to first token separately from generation speed, then compare a short prompt with a representative long one. Check whether the application actually reuses prefixes and whether memory pressure rises. Also measure tool time and successful task completion: this document-summary experiment does not explain a particular agent's repeated compaction or long end-to-end wait.
Continue with the local LLM latency checklist, the Qwen3.8 model page and separate RTX 5090 evidence, or the current Mac model download shortlist.