What hardware do you need to run GLM-5.3-Flash?
GLM-5.3-Flash at UD-Q4_K_M is 189 GB of weights, with a context ceiling of 1024k tokens. Not yet rated: released after our last ratings pass. Ties Qwen3.8 Flash for the strongest open model, but needs a 256 GB machine.
How good is it, really?
On the Artificial Analysis Intelligence Index v4.3 it scores 42, which puts it in the Sonnet-class band. In the same league as the labs' mainstream models on this index. Score source. See the whole table.
- Summarising — not rated
- Translation — not rated
- Everyday coding — not rated
- Reasoning & maths — not rated
- Agentic work — not rated
What it costs either way
Renting the same model costs $0.075 per million input tokens and $0.25 per million output (OpenRouter, cheapest active endpoint, checked 2026-09-09). Buying a machine only beats that if you use it hard enough, for long enough, that the hardware price divides down below the rental bill.
Machines that run it
| Machine | Price | Speed at 32k | Pay-back | |
|---|---|---|---|---|
| Mac Studio M5 Ultra, 256GB | $10,799 | 32 tok/s estimated | Pays back in 965 years | Run the numbers |
One machine per family, cheapest first. Speeds are measured where a public benchmark exists and estimated from memory bandwidth otherwise; the calculator says which for any configuration.
The specifics
- Parameters
- 321.3B, of which 18B are active per token
- Quantisation
- UD-Q4_K_M
- Weights on disk
- 189 GB
- KV cache
- 0.6 GB at 32k context — 34 of 45 layers use linear attention with a fixed state; the other 11 use latent attention with no RoPE part (1,024 B) plus a sparse-attention indexer (514 B). Note this is the figure for vLLM or SGLang: the reference Transformers code expands the latent before caching and uses far more.
- Maximum context
- 1024k tokens (1M)
- Licence
- MIT
- Sources
- source 1