Sunk Cost sunkcost.ai Data checked 2026-09-03

Mistral Small 3.2 24B Instruct vs Devstral Small 2 24B

Devstral Small 2 24B scores higher on the intelligence index, 8 against 7. Both take the same machine to start: the cheapest here that runs either is the Framework Desktop, 32GB, at $1,269. On it, they run at much the same speed: 9.7 and 9.7 tok/s, both estimated from memory bandwidth. At 500k tokens a day the Framework Desktop, 32GB pays for itself in 163 years running Mistral Small 3.2 24B Instruct, and never running Devstral Small 2 24B.

Mistral Small 3.2 24B InstructDevstral Small 2 24B
Intelligence index78
ClassBelow every hosted tierBelow every hosted tier
Weights14 GB14 GB
Needs at 32k20 GB20 GB
QuantisationQ4_K_MQ4_K_M
Parameters24B24B
Max context128k384k
API price per 1M$0.075 in / $0.2 out$0.02 in / $0.1 outpriced as gpt-oss-20b
LicenceApache 2.0Apache 2.0
Machines here that run it30 of 3730 of 37
Cheapest machine that runs itStrix Halo Framework Desktop, 32GB $1,269Strix Halo Framework Desktop, 32GB $1,269
Summarising good not rated
Translation good not rated
Everyday coding usable not rated
Reasoning & maths usable not rated
Agentic work usable not rated

Run Mistral Small 3.2 24B Instruct on the Framework Desktop, 32GB · or Devstral Small 2 24B

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

What the newer model changes

Mistral Small 3.2 24B Instruct is last generation. Devstral Small 2 24B is the current dense Mistral model of the same size, 24B. On the intelligence index it scores 8 where Mistral Small 3.2 24B Instruct scores 7. Both ask the same of the machine: 20 GB at 32k of context, the same weights and the same key-value cache. The same 30 of the 37 machines priced here run both, from the Framework Desktop, 32GB at $1,269.

Devstral Small 2 24B takes 384k of context where Mistral Small 3.2 24B Instruct stops at 128k, which is a ceiling rather than a setting: what you actually get is whatever the machine has room for.

Side by side on the Framework Desktop, 32GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Strix Halo Framework Desktop, 32GB, at $1,269.

Mistral Small 3.2 24B InstructDevstral Small 2 24B
Speed at 32k9.7 tok/s estimated9.7 tok/s estimated
Pay-back on this machinePays back in 163 yearsNever pays back
API cost per month$1.26$0.38priced as gpt-oss-20b

Run Mistral Small 3.2 24B Instruct on the Framework Desktop, 32GB · or Devstral Small 2 24B

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Framework Desktop, 32GB. Mistral Small 3.2 24B Instruct pays for it sooner at every level of use: Devstral Small 2 24B never pays for it at all.

A day's useMistral Small 3.2 24B InstructDevstral Small 2 24B
50ka few chats a day1,633 yearsNever pays back
200klight assistant use408 yearsNever pays back
1Ma moderate coding-assistant day82 yearsNever pays back
4Mheavy coding with an agent20 yearsNever pays back
20Magents running most of the day6.1 yearsits ceilingNever pays backits ceiling

The Framework Desktop, 32GB cannot generate 20M tokens a day on either model: at most 13.5M on Mistral Small 3.2 24B Instruct and 13.5M on Devstral Small 2 24B. Both figures on that row are for the most it can do.

Run Mistral Small 3.2 24B Instruct at 20M tokens a day · or Devstral Small 2 24B

Memory is not what separates them

Mistral Small 3.2 24B Instruct needs 20 GB of memory at 32k of context and Devstral Small 2 24B needs 20 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

Mistral Small 3.2 24B Instruct is also head to head with Qwen3 32B above it on the leaderboard and Qwen3 14B below it. Devstral Small 2 24B is also head to head with LFM2.5 2.6B above it on the leaderboard and Llama 3.1 8B Instruct below it.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Mistral Small 3.2 24B Instruct · every machine that runs Devstral Small 2 24B · every other match-up · both against the frontier · the quickest pay-back at each level of use