Best models for a 512GB Mac Studio: speed and memory
Published 2026-07-03 · updated 2026-09-23
Looking for a measured smaller-Mac setup? Our Qwen3.8-27B test on a 48 GiB M5 Max covers input lengths up to 12,263 tokens, first-token latency and prefix preparation. It is separate from the 512 GB download planning below.
For a 512 GB Mac Studio, start with a shortlist based on your task: a coding model for your repository, a general model for your prompts, and a memory budget for your intended context length. The current download shortlist below identifies exact MLX conversions and their file sizes. The separate historical M3 Ultra measurements show speed and first-token trade-offs; neither list establishes a quality winner.
M5 Ultra: 512 GB availability and benchmark evidence
As of September 23, 2026, Apple says the 512 GB M5 Ultra Mac Studio is coming in late October. The September 22 launch announcement confirms that the new Mac Studio range is available, but gives that later date for the 512 GB configuration. Check the delivery estimate for your exact configuration and region before planning a project around it.
Apple's technical specifications list 96 GB for the base M5 Ultra and 256 GB or 512 GB options with the 80-core GPU, plus 1.2 TB/s memory bandwidth. These are hardware specifications, not measured model speeds or usable context limits. We have no M5 Ultra measurement in this guide: the 512 GB results below are an external M3 Ultra report, and our linked 48 GiB experiment uses an M5 Max.
What to check before choosing M5 Ultra over your current machine
- The exact configuration: confirm the chip, GPU cores and installed memory. A result from a 96 GB or 256 GB review unit does not test a model that needs a 512 GB machine.
- Your actual request: look for the model file and precision, runtime/version, input and output lengths, and whether the prompt was cached. Compare first-token wait as well as generation speed at your intended context.
- Your actual task: check useful answers or completed coding tasks. Text-generation results do not establish image or video generation performance; those need the exact workflow, resolution and duration tested.
If your current machine can run the chosen model, record that baseline first. If it cannot, establish the complete memory requirement before selecting a tier. The memory checklist below and pinned downloads help you narrow that decision without treating a specification or a different chip's result as a benchmark of your planned setup.
Current MLX model downloads for a 512 GB Mac
Start with one 27B conversion to check your app and prompts, then try a larger model only if it improves a task you care about. These are community conversions of current Qwen and GLM models, checked September 12, 2026. We verified repository metadata and instructions. We subsequently tested the listed Qwen3.8-27B 4-bit revision on a 48 GiB M5 Max on September 16; the other three conversions have not been run by us on a Mac. None of these four has been tested by us on a 512 GB Mac. Their links identify fixed revisions so you can inspect the same files.
| Exact conversion / files | Weight files | When to consider it |
|---|---|---|
| Qwen3.8-27B · 4-bit | 16.05 GB 14.95 GiB · 3 shards | Smaller first download to establish a working baseline. Conversion used mlx-vlm 0.6.8. Text-only measurements on a 48 GiB M5 Max below. |
| Qwen3.8-27B · 8-bit | 29.50 GB 27.48 GiB · 6 shards | Higher-precision alternative for the same prompts. Conversion used mlx-vlm 0.6.8; extra bytes do not prove better answers. |
| Qwen3.8-Flash-Next · 4-bit | 111.52 GB 103.86 GiB · 22 shards | Larger-capacity candidate. Use the corrected conversion and supported mlx-vlm revision described below. |
| GLM-5.3 · 4-bit | 418.32 GB 389.59 GiB · 91 shards | Advanced candidate: its card names a patched mlx-lm build. Resolve that runtime requirement before a large download. |
Download all weight shards and the accompanying configuration/tokenizer files from the chosen repository; one shard is not the full model. Keep the MLX conversion with an MLX-compatible app. The GGUF files in our RTX 5090 study are different artifacts and do not establish performance for these MLX versions. Download the checked file inventory for exact byte counts, revisions and configuration sources.
Check the runtime before downloading 112–418 GB
The Flash-Next conversion card describes a normalization error in some earlier conversions that produced incoherent output. This revision was converted using mlx-vlm after its Qwen support change, at commit d1bd74ed. Follow the card's runtime instructions and verify a short coherent response before a long-context test. A similarly named file or a generic MLX integration is not proof that the fix is present.
The GLM conversion card targets a 512 GB M3 Ultra and names mlx-lm 0.31.3 with PR #1410. That change is still a draft in the upstream repository as checked September 12. Do not assume a standard install includes it. The author's fit claim is not our measurement: 418.32 GB of weight files leaves substantial runtime and context questions. Confirm the exact working build and usable memory before choosing it.
How much context will these downloads support?
The pinned Qwen configurations list 262,144 positions; the GLM configuration lists 1,048,576. These are configuration ceilings, not measured usable context on your Mac. File size does not account for the context cache, temporary allocations, macOS or other apps, and the models use different attention designs. Subtracting download size from 512 GB cannot produce a reliable token limit.
- Load one model and confirm useful output with a short prompt. Record the app/runtime version and model revision.
- Try one request at roughly 4k input tokens, then 16k, then your actual target length. Keep the output limit and task comparable.
- Record peak memory, memory pressure/swap, first-token wait and answer usefulness. Stop increasing context if the session becomes unstable or unacceptably slow.
This is a validation sequence, not a claimed capacity result. Report the longest successfully tested workload alongside its settings. Our historical 96 GB runs below cannot fill the current-model, 512 GB measurement gap.
Does Qwen3.8-27B need a 512 GB Mac?
The listed 4-bit revision completed our text-generation test on a 48 GiB M5 Max. That answers one narrow memory-fit question, but fitting the model did not make long prompts quick. In our September 16, 2026 study, median first-token wait rose from 8.17 seconds at 840 input tokens to 113.21 seconds at 12,263 tokens. Median stream decode stayed around 13–14 tok/s.
| Actual input tokens | Median first token | Observed range |
|---|---|---|
| 840 | 8.17 s | 3.79–11.78 s |
| 3,156 | 27.66 s | 16.59–35.35 s |
| 12,263 | 113.21 s | 74.39–133.19 s |
The exact artifact was mlx-community/Qwen3.8-27B-4bit, revision 3e6447f, with MLX 0.32.1 and mlx-lm 0.31.3 on a 40-GPU-core M5 Max. These were synthetic document-summary prompts, batch one, thinking off, with the model already loaded. Every response reached the 256-token cap; we did not score finished summaries or coding quality. This exploratory study is separate from signed suite-v1 leaderboard submissions.
Background load was uncontrolled and timings varied substantially. An earlier smoke attempt was stopped after system swap grew; the successful study does not establish a minimum memory tier or zero paging. It also does not predict M3 Ultra or M5 Ultra speed, maximum context, image input or concurrent users. For a Mac you already own, test your actual input length and memory pressure before using these results to decide on an upgrade.
Read the complete study, memory observations and reproduction method or inspect the raw prompts, outputs and timing events and all 12 measurements as CSV. The download also includes three prepared-prefix requests; their timing excludes prefix preparation, so they are not mixed into this fresh-cache table.
Published measurements on a 512 GB M3 Ultra
AI KIZAI reports these July 19–22, 2026 results from its 80-core GPU, 512 GB M3 Ultra. It also rents that hardware, so this is vendor-published evidence. We have not independently reproduced it or received these results as signed llm-speed submissions.
| Model / format | Input tokens | Generation tok/s | PP tok/s |
|---|---|---|---|
| DeepSeek R1-0528 · 4-bit MLX | 2,777 | 20.3 | 206.9 |
| gpt-oss-120B · MXFP4/BF16 MLX | 2,719 | 79.2 | 1,405.9 |
| Qwen3-Coder-Next · 4-bit MLX | 2,812 | 77.2 | 2,099.8 |
The publisher's JSON summary names macOS 26.5.1, MLX 0.32.0 and mlx-lm 0.31.3. It does not expose the per-artifact revision hashes or individual repetition logs. Use it as an attributed reference, not an exact reproduction package, current-model ranking or matched 96 GB versus 512 GB comparison. These July models differ from the September downloads above. Speed alone does not establish answer quality.
Historical signed submissions on a 96 GB M3 Ultra
These submissions report an M3 Ultra with a 60-core GPU and 96 GB of unified memory, not a 512 GB configuration. They use MLX 0.31.3 and 4-bit model variants on the suite-v1 chat-short workload. Treat them as historical reference points, not a forecast for every Mac Studio or a current ranking of all available models. Each speed links to its source run.
| Model (4-bit) | Decode tok/s | Time to first token |
|---|---|---|
| Qwen3-Coder-30B-A3B | 112.2 | 0.539 s |
| Qwen3-Next-80B-A3B | 80.34 | 4.493 s |
| Codestral-22B-v0.1 | 47.49 | 0.559 s |
| Llama-3.3-70B-Instruct | 16.78 | 5.420 s |
Qwen3-Next streams faster than Codestral in these runs, but takes longer to produce its first token. That distinction matters for short interactions. These are separate submissions, not repeated trials controlling every condition; model choice, caching and runtime settings can affect the result. See decode, prefill and time to first token before comparing a single headline number. A signature identifies a submission; it does not independently verify the reported hardware or model quality.
Does 512 GB make a model faster?
More memory can let you load larger models, retain more context or avoid swapping. It does not by itself guarantee faster decoding for a model that already fits. Check the GPU configuration, backend and workload as well as memory capacity. We do not have a matched 96 GB versus 512 GB test in the runs above.
Allow room for model weights, the context cache, runtime allocations, macOS and other applications. A model that loads at a short context can still exceed the available memory during a longer session. Backend support is another requirement: enough memory does not guarantee that your inference engine can run a model.
Planning a main model plus smaller helpers for a local agent? Use the multiple-model residency and queue checklist before choosing a memory tier. It separates model switching from overlapping requests; a single-model speed result does not test that combined workload.
Before paying for 512 GB: check the model and context together
Write down the exact model file, quantization, inference app and context length you plan to use. Compare that complete workload on the memory tiers you are considering. A larger memory pool is useful when it enables a workload you need; it is not a reason to assume a faster answer or better model quality.
- Check the file and runtime first. Verify the selected artifact's download size and supported backend. A planned offload feature is not a measured working configuration; ask for a completed run using the version you can install.
- Budget for the session, not only the weights. Include context-cache memory, runtime allocations, macOS and other applications. Concurrent requests can also change the memory requirement. Test the longest input you expect to use.
- Compare the experience you need. For coding, check correct task completion. For documents, check the first-token wait and answer usefulness at the intended length. Short-chat decode speed alone cannot answer either question.
Check a model's estimated memory fit to make an initial shortlist. Choose the model, quantization, context and hardware. The checker uses coarse estimates and the catalogued hardware capacity, which may represent the maximum memory configuration. Confirm it matches your machine; a fit estimate is not a tested context limit or proof of backend support.
Need measured evidence next? Inspect the submitted M3 Ultra runs and check each run's actual memory and GPU configuration. For the separate capacity-versus-speed decision, use the Mac Studio and RTX 5090 comparison.
Choose for your workload
- Coding: try the Qwen3-Coder configuration above on representative repository tasks, checking correctness and time to first token as well as output speed.
- General use: compare Qwen3-Next with a smaller model on your actual prompts. A larger total parameter count is not proof of better answers.
- Long documents: test at your intended context length and watch memory use. The short-chat table cannot establish long-context performance.
- A new hardware purchase: establish what your chosen model needs before paying for additional memory. These four configurations were reported on a 96 GB machine.
Explore the submitted configurations on the M3 Ultra hardware page or compare the hardware trade-offs in RTX 5090 vs M3 Ultra.
Help fill the 512 GB measurement gap
If you have a 512 GB Mac Studio, a comparable run would add evidence this guide currently lacks. Start with the benchmark methodology and the linked run's model, quantization, backend and workload. The command below starts the standard benchmark flow; it does not reproduce every model configuration automatically.
$ pipx install llm-speed && llm-speed benchM5 Ultra availability and buying guidance checked September 23, 2026. Model download inventory checked September 12. External 512 GB report checked September 15; signed source run pages reviewed September 5. Measurements were submitted April 28 and May 12, 2026; this article update is not a new benchmark run. Consult the cheatsheet for other submitted configurations and check their dates.