Sunk Cost sunkcost.ai Data checked 2026-09-03

Devstral Small 2 24B vs Gemma 3 27B it

Devstral Small 2 24B scores higher on the intelligence index, 8 against 5. Both take the same machine to start: the cheapest here that runs either is the Framework Desktop, 32GB, at $1,269. On it, they run at much the same speed: 9.7 and 9.8 tok/s, both estimated from memory bandwidth. At 500k tokens a day the Framework Desktop, 32GB pays for itself in 110 years running Gemma 3 27B it, and never running Devstral Small 2 24B.

Devstral Small 2 24BGemma 3 27B it
Intelligence index85
ClassBelow every hosted tierBelow every hosted tier
Weights14 GB17 GB
Needs at 32k20 GB20 GB
QuantisationQ4_K_MQ4_K_M
Parameters24B27.4B
Max context384k128k
API price per 1M$0.02 in / $0.1 outpriced as gpt-oss-20b$0.08 in / $0.45 out
LicenceApache 2.0Gemma Terms of Use
Machines here that run it30 of 3730 of 37
Cheapest machine that runs itStrix Halo Framework Desktop, 32GB $1,269Strix Halo Framework Desktop, 32GB $1,269
Summarising not rated good
Translation not rated good
Everyday coding not rated usable
Reasoning & maths not rated usable
Agentic work not rated don’t

Run Devstral Small 2 24B on the Framework Desktop, 32GB · or Gemma 3 27B it

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

Side by side on the Framework Desktop, 32GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Strix Halo Framework Desktop, 32GB, at $1,269.

Devstral Small 2 24BGemma 3 27B it
Speed at 32k9.7 tok/s estimated9.8 tok/s estimated
Pay-back on this machineNever pays backPays back in 110 years
API cost per month$0.38priced as gpt-oss-20b$1.57

Run Devstral Small 2 24B on the Framework Desktop, 32GB · or Gemma 3 27B it

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Framework Desktop, 32GB. Gemma 3 27B it pays for it sooner at every level of use: Devstral Small 2 24B never pays for it at all.

A day's useDevstral Small 2 24BGemma 3 27B it
50ka few chats a dayNever pays back1,105 years
200klight assistant useNever pays back276 years
1Ma moderate coding-assistant dayNever pays back55 years
4Mheavy coding with an agentNever pays back14 years
20Magents running most of the dayNever pays backits ceiling4.1 yearsits ceiling

The Framework Desktop, 32GB cannot generate 20M tokens a day on either model: at most 13.5M on Devstral Small 2 24B and 13.5M on Gemma 3 27B it. Both figures on that row are for the most it can do.

Run Devstral Small 2 24B at 20M tokens a day · or Gemma 3 27B it

Memory is not what separates them

Devstral Small 2 24B needs 20 GB of memory at 32k of context and Gemma 3 27B it needs 20 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

Devstral Small 2 24B is also head to head with LFM2.5 2.6B above it on the leaderboard, Llama 3.1 8B Instruct below it and Mistral Small 3.2 24B Instruct, the last-generation Mistral nearest it in size. One more model needs much the same memory: GLM-4.7-Flash. Gemma 3 27B it is also head to head with Qwen3 8B above it on the leaderboard, Ministral 3 8B below it and Gemma 4 31B it, the current Gemma nearest it in size. Two more models need much the same memory: Qwen3.6 27B and Mistral Small 3.2 24B Instruct.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Devstral Small 2 24B · every machine that runs Gemma 3 27B it · every other match-up · both against the frontier · the quickest pay-back at each level of use