Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run Ornith 1.5 35B-A3B?

Ornith 1.5 35B-A3B at Q4_K_M is 22 GB of weights, with a context ceiling of 256k tokens. Not yet rated: released after our last ratings pass. MIT-licensed 36B mixture with 3B active. Active-parameter count is inferred from the name, not stated on the card.

Cheapest machine that runs itStrix Halo Framework Desktop, 32GB at $1,269
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 145 tok/s at 32k context
Honest answer on costPays back in 83 years at 500k tokens a day.

How good is it, really?

It has not been placed on the intelligence index yet. See the whole table.

What it costs either way

Nobody rents Ornith 1.5 35B-A3B by the token. The closest hosted match, Qwen3.6 35B-A3B, costs $0.05 per million input tokens and $0.7 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-03). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 32GB $1,269 57 tok/s estimated Pays back in 83 years Run the numbers
Mac mini M5 Pro, 48GB $2,299 37 tok/s estimated Pays back in 158 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 56 tok/s estimated Pays back in 165 years Run the numbers
MacBook Pro M5 Pro (16-inch), 48GB $3,599 37 tok/s estimated Pays back in 248 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 71 tok/s estimated Pays back in 305 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
36B, of which 3B are active per token
Quantisation
Q4_K_M
Weights on disk
22 GB
KV cache
0.7 GB at 32k context — Hybrid: three linear-attention layers per full-attention layer. Multimodal; the vision projector is a separate 0.90 GB file, not counted here.
Maximum context
256k tokens (256k)
Licence
MIT
Sources
source 1, source 2