Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a Mac mini M4, 32GB run local LLMs?

Yes — 19 of the 39 open models on this site fit in its 21 GB of usable memory, the strongest being Qwen3.8 27B. Whether that saves you money is a different question, and the answer is usually no.

Price$999 at launch — discontinued
Memory32 GB, about 21 GB of it addressable by the GPU at 120 GB/s
Best model it runsQwen3.8 27B — Sonnet-class, 5 tok/s
Pay-back against the APIPays back in 13 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
Qwen3.8 27BQ4_K_M 4.8 tok/s Sonnet-class 19 GB
Qwen3.6 27BQ4_K_M 4.7 tok/s Haiku-class 19 GB
Muse Glimmer 30BQ4_K_M 5 tok/s Haiku-class 18 GB
Gemma 4 26B-A4BQ4_K_M 10 tok/s Haiku-class 18 GB
GLM-4.7-FlashQ4_K_M 10 tok/s Haiku-class 20 GB
Gemma 4 12BQ4_K_M 11 tok/s Below every hosted tier 8.0 GB
Qwen3.5 9BQ4_K_M 13 tok/s Below every hosted tier 6.8 GB
MiniCPM5 2BQ4_K_M 30 tok/s Below every hosted tier 3.0 GB
Qwen3.5 4BQ4_K_M 24 tok/s Below every hosted tier 3.8 GB
Ling 3.0 tinyQ4_K_M 15 tok/s Below every hosted tier 6.4 GB
Granite 4.2 8BQ4_K_M 8.4 tok/s Below every hosted tier 11 GB
gpt-oss-20bMXFP4 12 tok/s Below every hosted tier 13 GB

7 more fit; the calculator lists them all.

The specifics

Chip
M4 — 10-core CPU / 10-core GPU
Memory bandwidth
120 GB/s
Usable by the GPU
21 GB — macOS lets the GPU wire roughly two-thirds of unified memory by default (llama.cpp discussion #2182: 66.7% at 32 GiB or less, 75% above). Raise it with `sudo sysctl iogpu.wired_limit_mb`. Memory is treated as GB throughout, which is slightly conservative.
Power under load
65 W (published) — Apple's published 'max' figure for this chip (support article 103253), a CPU+GPU stress number; inference typically draws less.
Availability
Discontinued 2026-08-25. Launch price; Apple no longer sells it new, so treat as a refurbished/used reference.
Sources
source 1, source 2, source 3