Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Devstral Small 2 24B?

Devstral Small 2 24B at Q4_K_M is 14 GB of weights, with a context ceiling of 384k tokens. Not yet rated: released after our last ratings pass. Mistral's coding model, with an unusually long context for its size.

Cheapest machine that runs itStrix Halo Framework Desktop, 32GB at $1,269
Shortest pay-backMac Studio M5 Max, 48GB — Pays back in 2,535 years
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 46 tok/s at 32k context
Honest answer on costAgainst the API, buying hardware for this model never pays for itself at ordinary usage.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 8, which puts it in the Below every hosted tier band. Fine for simple, well-specified tasks. Noticeably less capable than anything the big labs sell today. Score source. See the whole table.

What it costs either way

Nobody rents Devstral Small 2 24B by the token. The closest hosted match, gpt-oss-20b, costs $0.02 per million input tokens and $0.1 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 32GB $1,269 9.7 tok/s estimated Never pays back Run the numbers
Mac mini M6, 32GB $1,299 6.5 tok/s estimated Never pays back Run the numbers
MacBook Pro M5 (14-inch), 32GB $2,399 5.8 tok/s estimated Never pays back Run the numbers
Mac Studio M5 Max, 36GB $2,499 18 tok/s estimated Pays back in 24,221 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 10 tok/s estimated Never pays back Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
24B
Quantisation
Q4_K_M
Weights on disk
14 GB
KV cache
5.4 GB at 32k context — All 40 layers full attention; no sliding window.
Maximum context
384k tokens (384k)
Licence
Apache 2.0
Sources
source 1