Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a Mac mini M4 Pro, 24GB run local LLMs?

Yes — 13 of the 39 open models on this site fit in its 16 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.

Price$1,399 at launch — discontinued
Memory24 GB, about 16 GB of it addressable by the GPU at 273 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 26 tok/s
Pay-back against the APIPays back in 114 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
Gemma 4 12BQ4_K_M 26 tok/s Below every hosted tier 8.0 GB
Qwen3.5 9BQ4_K_M 30 tok/s Below every hosted tier 6.8 GB
MiniCPM5 2BQ4_K_M 69 tok/s Below every hosted tier 3.0 GB
Qwen3.5 4BQ4_K_M 54 tok/s Below every hosted tier 3.8 GB
Ling 3.0 tinyQ4_K_M 34 tok/s Below every hosted tier 6.4 GB
Granite 4.2 8BQ4_K_M 19 tok/s Below every hosted tier 11 GB
gpt-oss-20bMXFP4 28 tok/s Below every hosted tier 13 GB
Gemma 4 E4BQAT Q4_0 24 tok/s Below every hosted tier 5.7 GB
LFM2.5 2.6BQ4_K_M 93 tok/s Below every hosted tier 2.2 GB
Ministral 3 14BQ4_K_M 15 tok/s Below every hosted tier 14 GB
Ministral 3 8BQ4_K_M 21 tok/s Below every hosted tier 9.8 GB
Ornith 1.5 9BQ4_K_M 30 tok/s not yet placed 6.9 GB

1 more fit; the calculator lists them all.

The specifics

Chip
M4 Pro — 12-core CPU / 16-core GPU
Memory bandwidth
273 GB/s
Usable by the GPU
16 GB — macOS lets the GPU wire roughly two-thirds of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. Memory is treated as GB throughout, which is slightly conservative.
Power under load
140 W (published) — Apple's published 'max' figure for this chip (support article 103253), a CPU+GPU stress number; inference typically draws less.
Availability
Discontinued 2026-08-25. Launch price; Apple no longer sells it new, so treat as a refurbished/used reference.
Sources
source 1, source 2, source 3