Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Laguna XS 2.1?

Laguna XS 2.1 at Q4_K_M is 20 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. poolside's agentic-coding MoE: 33B total but only 3B active per token.

Cheapest machine that runs itStrix Halo Framework Desktop, 32GB at $1,269
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 112 tok/s at 32k context
Honest answer on costPays back in 127 years at 500k tokens a day.

How good is it, really?

It has not been placed on the intelligence index yet. See the whole table.

What it costs either way

Renting the same model costs $0.06 per million input tokens and $0.12 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 32GB $1,269 44 tok/s estimated Pays back in 127 years Run the numbers
Mac mini M5 Pro, 48GB $2,299 29 tok/s estimated Pays back in 255 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 43 tok/s estimated Pays back in 255 years Run the numbers
MacBook Pro M5 Pro (16-inch), 48GB $3,599 29 tok/s estimated Pays back in 400 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 55 tok/s estimated Pays back in 462 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
33.4B, of which 3B are active per token
Quantisation
Q4_K_M
Weights on disk
20 GB
KV cache
1.4 GB at 32k context — 3 sliding-window layers (512) per full-attention layer. The runtime keeps the KV cache in FP8, so real usage is about half the figure here.
Maximum context
256k tokens (256k)
Licence
OpenMDW 1.1
Sources
source 1, source 2