What hardware do you need to run Laguna XS 2.1?
Laguna XS 2.1 at Q4_K_M is 20 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. poolside's agentic-coding MoE: 33B total but only 3B active per token.
How good is it, really?
It has not been placed on the intelligence index yet. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Renting the same model costs $0.06 per million input tokens and $0.12 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Strix Halo Framework Desktop, 32GB | $1,269 | 44 tok/s estimated | Pays back in 127 years | Run the numbers |
| Mac mini M5 Pro, 48GB | $2,299 | 29 tok/s estimated | Pays back in 255 years | Run the numbers |
| Mac Studio M5 Max, 36GB | $2,499 | 43 tok/s estimated | Pays back in 255 years | Run the numbers |
| MacBook Pro M5 Pro (16-inch), 48GB | $3,599 | 29 tok/s estimated | Pays back in 400 years | Run the numbers |
| DGX Spark GB10 Grace Blackwell, 128GB | $4,699 | 55 tok/s estimated | Pays back in 462 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 33.4B, of which 3B are active per token
- Quantisation
- Q4_K_M
- Weights on disk
- 20 GB
- KV cache
- 1.4 GB at 32k context — 3 sliding-window layers (512) per full-attention layer. The runtime keeps the KV cache in FP8, so real usage is about half the figure here.
- Maximum context
- 256k tokens (256k)
- Licence
- OpenMDW 1.1
- Sources
- source 1, source 2