Can a MacBook Pro M5 Pro (16-inch), 64GB run local LLMs?
Yes — 27 of the 39 open models on this site fit in its 48 GB of usable memory, the strongest being Qwen3.8 27B. Whether that saves you money is a different question, and the answer is usually no.
Price$3,999
Memory64 GB, about 48 GB of it addressable by the GPU at 307 GB/s
Best model it runsQwen3.8 27B — Sonnet-class, 12 tok/s
Pay-back against the APIPays back in 52 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Qwen3.8 27BQ4_K_M | 12 tok/s | Sonnet-class | 19 GB | |
| Qwen3.6 27BQ4_K_M | 12 tok/s | Haiku-class | 19 GB | |
| Qwen3.6 35B-A3BQ4_K_M | 37 tok/s | Haiku-class | 23 GB | |
| Muse Glimmer 30BQ4_K_M | 13 tok/s | Haiku-class | 18 GB | |
| Gemma 4 26B-A4BQ4_K_M | 27 tok/s | Haiku-class | 18 GB | |
| Gemma 4 31B itQ4_K_M | 8.9 tok/s | Haiku-class | 26 GB | |
| Granite 4.2 30BQ4_K_M | 8.8 tok/s | Haiku-class | 26 GB | |
| GLM-4.7-FlashQ4_K_M | 26 tok/s | Haiku-class | 20 GB | |
| Gemma 4 12BQ4_K_M | 29 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 34 tok/s | Below every hosted tier | 6.8 GB | |
| Nemotron 3.5 Lightning 30B-A3BQ4_K_M | 35 tok/s | Below every hosted tier | 26 GB | |
| MiniCPM5 2BQ4_K_M | 78 tok/s | Below every hosted tier | 3.0 GB |
15 more fit; the calculator lists them all.
The specifics
- Chip
- M5 Pro (16-inch) — 15-core CPU / 16-core GPU
- Memory bandwidth
- 307 GB/s
- Usable by the GPU
- 48 GB — macOS lets the GPU wire roughly 75% of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. On a laptop, sustained speed drops once the chassis warms up and the fans cap out, so a desktop with the same chip will hold a higher tokens/sec over a long generation than the figures here suggest.
- Power under load
- 140 W (stand in) — Apple has published no power figures for the 2026 Macs, and none at all for sustained GPU load on a laptop. Stand-in: the desktop maximum for the equivalent chip (M4 Pro Mac mini, 140 W). A laptop draws less than that, so the electricity line here is an over-estimate.
- Availability
- Shipping.
- Sources
- source 1, source 2, source 3