Sunk Cost sunkcost.ai Data checked 2026-09-03

Gemma 3 12B it vs Gemma 4 12B

Gemma 4 12B scores higher on the intelligence index, 14 against 4. Both take the same machine to start: the cheapest here that runs either is the Mac mini M6, 16GB, at $899. On it, Gemma 4 12B is about 1.2× quicker: 14 tok/s against 12, both estimated from memory bandwidth. At 500k tokens a day the Mac mini M6, 16GB pays for itself in 71 years running Gemma 4 12B, against 123 years running Gemma 3 12B it.

Gemma 3 12B itGemma 4 12B
Intelligence index414
ClassBelow every hosted tierBelow every hosted tier
Weights7.3 GB7.1 GB
Needs at 32k9.8 GB8.0 GB
QuantisationQ4_K_MQ4_K_M
Parameters12.2B12B
Max context128k256k
API price per 1M$0.05 in / $0.15 out$0.08 in / $0.13 outpriced as Qwen3.5 9B
LicenceGemma Terms of UseApache 2.0
Machines here that run it37 of 3737 of 37
Cheapest machine that runs itMac mini M6, 16GB $899Mac mini M6, 16GB $899
Summarising good not rated
Translation good not rated
Everyday coding usable not rated
Reasoning & maths don’t not rated
Agentic work don’t not rated

Run Gemma 3 12B it on the Mac mini M6, 16GB · or Gemma 4 12B

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

What the newer model changes

Gemma 3 12B it is last generation. Gemma 4 12B is the current dense Gemma model nearest it in size, 12B against 12.2B. On the intelligence index it scores 14 where Gemma 3 12B it scores 4. It asks less of the machine: 8.0 GB at 32k of context against 9.8 GB. The weights are 7.1 GB against 7.3 GB, and the cache at that window is 0.9 GB against 2.5 GB. Every one of the 37 machines priced here runs both, from the Mac mini M6, 16GB at $899.

Gemma 4 12B takes 256k of context where Gemma 3 12B it stops at 128k, which is a ceiling rather than a setting: what you actually get is whatever the machine has room for.

Side by side on the Mac mini M6, 16GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Mac mini M6, 16GB, at $899.

Gemma 3 12B itGemma 4 12B
Speed at 32k12 tok/s estimated14 tok/s estimated
Pay-back on this machinePays back in 123 yearsPays back in 71 years
API cost per month$0.86$1.27priced as Qwen3.5 9B

Run Gemma 3 12B it on the Mac mini M6, 16GB · or Gemma 4 12B

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Mac mini M6, 16GB. Gemma 4 12B pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.

A day's useGemma 3 12B itGemma 4 12B
50ka few chats a day1,234 years706 years
200klight assistant use308 years176 years
1Ma moderate coding-assistant day62 years35 years
4Mheavy coding with an agent15 years8.8 years
20Magents running most of the day3.8 yearsits ceiling21 monthsits ceiling

The Mac mini M6, 16GB cannot generate 20M tokens a day on either model: at most 16.2M on Gemma 3 12B it and 19.8M on Gemma 4 12B. Both figures on that row are for the most it can do.

Run Gemma 3 12B it at 20M tokens a day · or Gemma 4 12B

Memory is not what separates them

Gemma 3 12B it needs 9.8 GB of memory at 32k of context and Gemma 4 12B needs 8.0 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

Gemma 3 12B it is also head to head with Ministral 3 8B above it on the leaderboard. Gemma 4 12B is also head to head with GLM-4.7-Flash above it on the leaderboard and Qwen3.5 9B below it.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Gemma 3 12B it · every machine that runs Gemma 4 12B · every other match-up · both against the frontier · the quickest pay-back at each level of use