Sunk Cost sunkcost.ai Data checked 2026-09-03

Gemma 3 27B it vs Gemma 4 31B it

Gemma 4 31B it scores higher on the intelligence index, 15 against 5. The cheapest machine here that runs Gemma 4 31B it is the Radeon AI PRO R9700, 32GB, at $1,299. Gemma 3 27B it runs on the Framework Desktop, 32GB at $1,269, $30 less. The Radeon AI PRO R9700, 32GB is priced as the card alone, without the PC around it. On the Radeon AI PRO R9700, 32GB, the cheapest machine here that runs both, Gemma 3 27B it is about 1.3× quicker: 28 tok/s against 22, both estimated from memory bandwidth. At 500k tokens a day the Radeon AI PRO R9700, 32GB pays for itself in 99 years running Gemma 3 27B it, against 110 years running Gemma 4 31B it.

Gemma 3 27B itGemma 4 31B it
Intelligence index515
ClassBelow every hosted tierHaiku-class
Weights17 GB20 GB
Needs at 32k20 GB26 GB
QuantisationQ4_K_MQ4_K_M
Parameters27.4B31.3B
Max context128k256k
API price per 1M$0.08 in / $0.45 out$0.09 in / $0.34 out
LicenceGemma Terms of UseApache 2.0
Machines here that run it30 of 3727 of 37
Cheapest machine that runs itStrix Halo Framework Desktop, 32GB $1,269AMD Radeon AI PRO R9700, 32GB $1,299card only
Summarising good not rated
Translation good not rated
Everyday coding usable not rated
Reasoning & maths usable not rated
Agentic work don’t not rated

Run Gemma 3 27B it on the Framework Desktop, 32GB · or Gemma 4 31B it on the Radeon AI PRO R9700, 32GB

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

What the newer model changes

Gemma 3 27B it is last generation. Gemma 4 31B it is the current dense Gemma model nearest it in size, 31.3B against 27.4B. On the intelligence index it scores 15 where Gemma 3 27B it scores 5. It asks more of the machine: 26 GB at 32k of context against 20 GB. The weights are 20 GB against 17 GB, and the cache at that window is 6.2 GB against 3.1 GB. 27 of the 37 machines priced here run it, against 30 for Gemma 3 27B it. The cheapest that runs it is the Radeon AI PRO R9700, 32GB at $1,299, card only, where Gemma 3 27B it starts at the Framework Desktop, 32GB at $1,269.

Gemma 4 31B it takes 256k of context where Gemma 3 27B it stops at 128k, which is a ceiling rather than a setting: what you actually get is whatever the machine has room for.

Side by side on the Radeon AI PRO R9700, 32GB

The table above gives each model the cheapest machine that runs it, and those are two different machines, so nothing in it is a like-for-like race. The AMD Radeon AI PRO R9700, 32GB is the cheapest machine here that runs both, so this is the pair doing the same work on the same hardware.

Gemma 3 27B itGemma 4 31B it
Speed at 32k28 tok/s estimated22 tok/s estimated
Pay-back on this machinePays back in 99 yearsPays back in 110 years
API cost per month$1.57$1.61

Run Gemma 3 27B it on the Radeon AI PRO R9700, 32GB · or Gemma 4 31B it

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Radeon AI PRO R9700, 32GB. Gemma 3 27B it pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.

A day's useGemma 3 27B itGemma 4 31B it
50ka few chats a day990 years1,101 years
200klight assistant use248 years275 years
1Ma moderate coding-assistant day50 years55 years
4Mheavy coding with an agent12 years14 years
20Magents running most of the day2.5 years2.8 years

Run Gemma 3 27B it at 20M tokens a day · or Gemma 4 31B it

Machines that run one and not the other

Gemma 3 27B it needs 20 GB of memory at 32k of context and Gemma 4 31B it needs 26 GB. That puts Gemma 3 27B it on 3 of the 37 machines priced here that Gemma 4 31B it does not, starting at $1,269.

MachinePriceMemorySpeed on Gemma 3 27B itPay-back
Strix Halo Framework Desktop, 32GB$1,26932 GB9.8 tok/s estimatedPays back in 110 years
Mac mini M6, 32GB$1,29932 GB6.5 tok/s estimatedPays back in 97 years
MacBook Pro M5 (14-inch), 32GB$2,39932 GB5.8 tok/s estimatedPays back in 187 years

Gemma 3 27B it is also head to head with Qwen3 8B above it on the leaderboard and Ministral 3 8B below it. Gemma 4 31B it is also head to head with Qwen3.5 122B-A10B above it on the leaderboard and Granite 4.2 30B below it.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Gemma 3 27B it · every machine that runs Gemma 4 31B it · every other match-up · both against the frontier · the quickest pay-back at each level of use