Gemma 3 27B it vs Gemma 4 31B it
Gemma 4 31B it scores higher on the intelligence index, 15 against 5. The cheapest machine here that runs Gemma 4 31B it is the Radeon AI PRO R9700, 32GB, at $1,299. Gemma 3 27B it runs on the Framework Desktop, 32GB at $1,269, $30 less. The Radeon AI PRO R9700, 32GB is priced as the card alone, without the PC around it. On the Radeon AI PRO R9700, 32GB, the cheapest machine here that runs both, Gemma 3 27B it is about 1.3× quicker: 28 tok/s against 22, both estimated from memory bandwidth. At 500k tokens a day the Radeon AI PRO R9700, 32GB pays for itself in 99 years running Gemma 3 27B it, against 110 years running Gemma 4 31B it.
| Gemma 3 27B it | Gemma 4 31B it | |
|---|---|---|
| Intelligence index | 5 | 15 |
| Class | Below every hosted tier | Haiku-class |
| Weights | 17 GB | 20 GB |
| Needs at 32k | 20 GB | 26 GB |
| Quantisation | Q4_K_M | Q4_K_M |
| Parameters | 27.4B | 31.3B |
| Max context | 128k | 256k |
| API price per 1M | $0.08 in / $0.45 out | $0.09 in / $0.34 out |
| Licence | Gemma Terms of Use | Apache 2.0 |
| Machines here that run it | 30 of 37 | 27 of 37 |
| Cheapest machine that runs it | Strix Halo Framework Desktop, 32GB $1,269 | AMD Radeon AI PRO R9700, 32GB $1,299card only |
| Summarising | good | not rated |
| Translation | good | not rated |
| Everyday coding | usable | not rated |
| Reasoning & maths | usable | not rated |
| Agentic work | don’t | not rated |
Run Gemma 3 27B it on the Framework Desktop, 32GB · or Gemma 4 31B it on the Radeon AI PRO R9700, 32GB
Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.
What the newer model changes
Gemma 3 27B it is last generation. Gemma 4 31B it is the current dense Gemma model nearest it in size, 31.3B against 27.4B. On the intelligence index it scores 15 where Gemma 3 27B it scores 5. It asks more of the machine: 26 GB at 32k of context against 20 GB. The weights are 20 GB against 17 GB, and the cache at that window is 6.2 GB against 3.1 GB. 27 of the 37 machines priced here run it, against 30 for Gemma 3 27B it. The cheapest that runs it is the Radeon AI PRO R9700, 32GB at $1,299, card only, where Gemma 3 27B it starts at the Framework Desktop, 32GB at $1,269.
Gemma 4 31B it takes 256k of context where Gemma 3 27B it stops at 128k, which is a ceiling rather than a setting: what you actually get is whatever the machine has room for.
Side by side on the Radeon AI PRO R9700, 32GB
The table above gives each model the cheapest machine that runs it, and those are two different machines, so nothing in it is a like-for-like race. The AMD Radeon AI PRO R9700, 32GB is the cheapest machine here that runs both, so this is the pair doing the same work on the same hardware.
| Gemma 3 27B it | Gemma 4 31B it | |
|---|---|---|
| Speed at 32k | 28 tok/s estimated | 22 tok/s estimated |
| Pay-back on this machine | Pays back in 99 years | Pays back in 110 years |
| API cost per month | $1.57 | $1.61 |
Run Gemma 3 27B it on the Radeon AI PRO R9700, 32GB · or Gemma 4 31B it
How much use it takes to pay for the machine
Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Radeon AI PRO R9700, 32GB. Gemma 3 27B it pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.
| A day's use | Gemma 3 27B it | Gemma 4 31B it |
|---|---|---|
| 50ka few chats a day | 990 years | 1,101 years |
| 200klight assistant use | 248 years | 275 years |
| 1Ma moderate coding-assistant day | 50 years | 55 years |
| 4Mheavy coding with an agent | 12 years | 14 years |
| 20Magents running most of the day | 2.5 years | 2.8 years |
Run Gemma 3 27B it at 20M tokens a day · or Gemma 4 31B it
Machines that run one and not the other
Gemma 3 27B it needs 20 GB of memory at 32k of context and Gemma 4 31B it needs 26 GB. That puts Gemma 3 27B it on 3 of the 37 machines priced here that Gemma 4 31B it does not, starting at $1,269.
| Machine | Price | Memory | Speed on Gemma 3 27B it | Pay-back |
|---|---|---|---|---|
| Strix Halo Framework Desktop, 32GB | $1,269 | 32 GB | 9.8 tok/s estimated | Pays back in 110 years |
| Mac mini M6, 32GB | $1,299 | 32 GB | 6.5 tok/s estimated | Pays back in 97 years |
| MacBook Pro M5 (14-inch), 32GB | $2,399 | 32 GB | 5.8 tok/s estimated | Pays back in 187 years |
Gemma 3 27B it is also head to head with Qwen3 8B above it on the leaderboard and Ministral 3 8B below it. Gemma 4 31B it is also head to head with Qwen3.5 122B-A10B above it on the leaderboard and Granite 4.2 30B below it.
The assumptions behind both columns
Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.
More head to head: every machine that runs Gemma 3 27B it · every machine that runs Gemma 4 31B it · every other match-up · both against the frontier · the quickest pay-back at each level of use