Can a Mac mini M5 Pro, 24GB run local LLMs?
Yes — 13 of the 39 open models on this site fit in its 16 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.
Price$1,699
Memory24 GB, about 16 GB of it addressable by the GPU at 307 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 29 tok/s
Pay-back against the APIPays back in 135 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Gemma 4 12BQ4_K_M | 29 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 34 tok/s | Below every hosted tier | 6.8 GB | |
| MiniCPM5 2BQ4_K_M | 78 tok/s | Below every hosted tier | 3.0 GB | |
| Qwen3.5 4BQ4_K_M | 60 tok/s | Below every hosted tier | 3.8 GB | |
| Ling 3.0 tinyQ4_K_M | 38 tok/s | Below every hosted tier | 6.4 GB | |
| Granite 4.2 8BQ4_K_M | 21 tok/s | Below every hosted tier | 11 GB | |
| gpt-oss-20bMXFP4 | 32 tok/s | Below every hosted tier | 13 GB | |
| Gemma 4 E4BQAT Q4_0 | 27 tok/s | Below every hosted tier | 5.7 GB | |
| LFM2.5 2.6BQ4_K_M | 104 tok/s | Below every hosted tier | 2.2 GB | |
| Ministral 3 14BQ4_K_M | 17 tok/s | Below every hosted tier | 14 GB | |
| Ministral 3 8BQ4_K_M | 24 tok/s | Below every hosted tier | 9.8 GB | |
| Ornith 1.5 9BQ4_K_M | 34 tok/s | not yet placed | 6.9 GB |
1 more fit; the calculator lists them all.
The specifics
- Chip
- M5 Pro — 15-core CPU / 16-core GPU
- Memory bandwidth
- 307 GB/s
- Usable by the GPU
- 16 GB — macOS lets the GPU wire roughly two-thirds of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. Memory is treated as GB throughout, which is slightly conservative.
- Power under load
- 140 W (stand in) — Apple has not published power figures for the 2026 machines yet (support articles 103253 / 102027 stop at the 2024/2025 models). Stand-in: Apple's published maximum for the previous chip (M4 Pro Mac mini, 140 W), which is a CPU+GPU stress figure, so it likely overstates inference draw.
- Availability
- Pre-order; ships 2026-09-22.
- Sources
- source 1, source 2, source 3