Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a Mac Studio M5 Ultra, 256GB run local LLMs?

Yes — 38 of the 39 open models on this site fit in its 192 GB of usable memory, the strongest being GLM-5.3-Flash. Whether that saves you money is a different question, and the answer is usually no.

Price$10,799
Memory256 GB, about 192 GB of it addressable by the GPU at 1200 GB/s
Best model it runsGLM-5.3-Flash — Sonnet-class, 32 tok/s
Pay-back against the APIPays back in 965 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
GLM-5.3-FlashUD-Q4_K_M 32 tok/s Sonnet-class 190 GB
Qwen3.8 Flash NextQ4_K_M 74 tok/s Sonnet-class 121 GB
DeepSeek V4-FlashUD-Q4_K_M 49 tok/s Sonnet-class 155 GB
Qwen3.8 27BQ4_K_M 48 tok/s Sonnet-class 19 GB
Inkling SmallUD-Q4_K_M 43 tok/s Haiku-class 164 GB
Ling 3.0 flashQ4_K_M 52 tok/s Haiku-class 82 GB
MiniMax M2.7Q4_K_M 25 tok/s Haiku-class 149 GB
Qwen3.6 27BQ4_K_M 47 tok/s Haiku-class 19 GB
Qwen3.6 35B-A3BQ4_K_M 143 tok/s Haiku-class 23 GB
Muse Glimmer 30BQ4_K_M 50 tok/s Haiku-class 18 GB
Gemma 4 26B-A4BQ4_K_M 105 tok/s Haiku-class 18 GB
Qwen3.5 122B-A10BUD-Q4_K_M 51 tok/s Haiku-class 79 GB

26 more fit; the calculator lists them all.

The specifics

Chip
M5 Ultra — 36-core CPU / 80-core GPU
Memory bandwidth
1200 GB/s
Usable by the GPU
192 GB — macOS lets the GPU wire roughly 75% of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. Memory is treated as GB throughout, which is slightly conservative.
Power under load
270 W (stand in) — Apple has not published power figures for the 2026 machines yet (support articles 103253 / 102027 stop at the 2024/2025 models). Stand-in: Apple's published maximum for the previous chip (M3 Ultra Mac Studio, 270 W), which is a CPU+GPU stress figure, so it likely overstates inference draw.
Availability
Pre-order; ships 2026-09-22.
Sources
source 1, source 2, source 3