Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a MacBook Pro M5 (14-inch), 16GB run local LLMs?

Yes — 10 of the 39 open models on this site fit in its 10.5 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.

Price$1,999
Memory16 GB, about 10.5 GB of it addressable by the GPU at 153 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 14 tok/s
Pay-back against the APIPays back in 157 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
Gemma 4 12BQ4_K_M 14 tok/s Below every hosted tier 8.0 GB
Qwen3.5 9BQ4_K_M 17 tok/s Below every hosted tier 6.8 GB
MiniCPM5 2BQ4_K_M 39 tok/s Below every hosted tier 3.0 GB
Qwen3.5 4BQ4_K_M 30 tok/s Below every hosted tier 3.8 GB
Ling 3.0 tinyQ4_K_M 19 tok/s Below every hosted tier 6.4 GB
Gemma 4 E4BQAT Q4_0 13 tok/s Below every hosted tier 5.7 GB
LFM2.5 2.6BQ4_K_M 52 tok/s Below every hosted tier 2.2 GB
Ministral 3 8BQ4_K_M 12 tok/s Below every hosted tier 9.8 GB
Ornith 1.5 9BQ4_K_M 17 tok/s not yet placed 6.9 GB
Spark-X2.5 4BQ4_K_M 30 tok/s not yet placed 3.9 GB

The specifics

Chip
M5 (14-inch) — 10-core CPU / 10-core GPU
Memory bandwidth
153 GB/s
Usable by the GPU
10.5 GB — macOS lets the GPU wire roughly two-thirds of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. On a laptop, sustained speed drops once the chassis warms up and the fans cap out, so a desktop with the same chip will hold a higher tokens/sec over a long generation than the figures here suggest.
Power under load
65 W (stand in) — Apple has published no power figures for the 2026 Macs, and none at all for sustained GPU load on a laptop. Stand-in: the desktop maximum for the equivalent chip (M4 Mac mini, 65 W). A laptop draws less than that, so the electricity line here is an over-estimate.
Availability
Shipping.
Sources
source 1, source 2, source 3