Can a Mac Studio M3 Ultra, 256GB run local LLMs?
Yes — 38 of the 39 open models on this site fit in its 192 GB of usable memory, the strongest being GLM-5.3-Flash. Whether that saves you money is a different question, and the answer is usually no.
Price$7,099 at launch — discontinued
Memory256 GB, about 192 GB of it addressable by the GPU at 819 GB/s
Best model it runsGLM-5.3-Flash — Sonnet-class, 22 tok/s
Pay-back against the APIPays back in 781 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| GLM-5.3-FlashUD-Q4_K_M | 22 tok/s | Sonnet-class | 190 GB | |
| Qwen3.8 Flash NextQ4_K_M | 50 tok/s | Sonnet-class | 121 GB | |
| DeepSeek V4-FlashUD-Q4_K_M | 34 tok/s | Sonnet-class | 155 GB | |
| Qwen3.8 27BQ4_K_M | 33 tok/s | Sonnet-class | 19 GB | |
| Inkling SmallUD-Q4_K_M | 29 tok/s | Haiku-class | 164 GB | |
| Ling 3.0 flashQ4_K_M | 35 tok/s | Haiku-class | 82 GB | |
| MiniMax M2.7Q4_K_M | 17 tok/s | Haiku-class | 149 GB | |
| Qwen3.6 27BQ4_K_M | 32 tok/s | Haiku-class | 19 GB | |
| Qwen3.6 35B-A3BQ4_K_M | 98 tok/s | Haiku-class | 23 GB | |
| Muse Glimmer 30BQ4_K_M | 34 tok/s | Haiku-class | 18 GB | |
| Gemma 4 26B-A4BQ4_K_M | 72 tok/s | Haiku-class | 18 GB | |
| Qwen3.5 122B-A10BUD-Q4_K_M | 35 tok/s | Haiku-class | 79 GB |
26 more fit; the calculator lists them all.
The specifics
- Chip
- M3 Ultra — 32-core CPU / 80-core GPU
- Memory bandwidth
- 819 GB/s
- Usable by the GPU
- 192 GB — macOS lets the GPU wire roughly 75% of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. Memory is treated as GB throughout, which is slightly conservative.
- Power under load
- 270 W (published) — Apple's published 'max' figure for this chip (support article 102027), a CPU+GPU stress number; inference typically draws less.
- Availability
- Discontinued 2026-08-25. Launch price; Apple no longer sells it new, so treat as a refurbished/used reference. 256GB required the 32-core chip.
- Sources
- source 1, source 2, source 3