Sunk Cost sunkcost.ai Data checked 2026-09-03

What hardware do you need to run GLM-4.7-Flash?

GLM-4.7-Flash at Q4_K_M is 18 GB of weights, with a context ceiling of 198k tokens. Not yet rated: released after our last ratings pass. Replaces GLM-4.5-Air at a quarter of the size and a higher score.

Cheapest machine that runs itStrix Halo Framework Desktop, 32GB at $1,269
Fastest of the ones listedMac Studio M5 Ultra, 96GB — 102 tok/s at 32k context
Honest answer on costPays back in 96 years at 500k tokens a day.

How good is it, really?

On the Artificial Analysis Intelligence Index v4.3 it scores 15 (reasoning mode; 11 without), which puts it in the Haiku-class band. In the same band as Anthropic's cheap, fast tier. Every current OpenAI model scores above this band. Score source. See the whole table.

What it costs either way

Renting the same model costs $0.061 per million input tokens and $0.4 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.

Machines that run it

MachinePriceSpeed at 32kPay-back
Strix Halo Framework Desktop, 32GB $1,269 40 tok/s estimated Pays back in 96 years Run the numbers
Mac mini M6, 32GB $1,299 14 tok/s estimated Pays back in 103 years Run the numbers
MacBook Pro M5 (14-inch), 32GB $2,399 13 tok/s estimated Pays back in 195 years Run the numbers
Mac Studio M5 Max, 36GB $2,499 39 tok/s estimated Pays back in 192 years Run the numbers
DGX Spark GB10 Grace Blackwell, 128GB $4,699 50 tok/s estimated Pays back in 351 years Run the numbers

One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.

The specifics

Parameters
31.2B, of which 3B are active per token
Quantisation
Q4_K_M
Weights on disk
18 GB
KV cache
1.8 GB at 32k context — Latent attention: the cache holds one 512-wide compressed vector plus 64 RoPE dimensions per layer, 1,152 bytes a token, not the 20 KB a head count would suggest. The config states no head_dim and it cannot be derived.
Maximum context
198k tokens (198k)
Licence
MIT
Sources
source 1