Sunk Cost sunkcost.ai Data checked 2026-09-03

Ling 3.0 flash vs GLM-4.5-Air

Ling 3.0 flash scores higher on the intelligence index, 25 against 11. Both take the same machine to start: the cheapest here that runs either is the Framework Desktop, 128GB, at $3,449. On it, Ling 3.0 flash is about 2× quicker: 20 tok/s against 10, both estimated from memory bandwidth. At 500k tokens a day the Framework Desktop, 128GB pays for itself in 139 years running GLM-4.5-Air, against 4,468 years running Ling 3.0 flash.

Ling 3.0 flashGLM-4.5-Air
Intelligence index2511
ClassHaiku-classBelow every hosted tier
Weights78 GB73 GB
Needs at 32k82 GB79 GB
QuantisationQ4_K_MQ4_K_M
Parameters124B (5.1B active)110.5B (12B active)
Max context256k128k
API price per 1M$0.021 in / $0.063 out$0.13 in / $0.85 out
LicenceMITMIT
Machines here that run it12 of 3712 of 37
Cheapest machine that runs itStrix Halo Framework Desktop, 128GB $3,449Strix Halo Framework Desktop, 128GB $3,449
Summarising not rated good
Translation not rated good
Everyday coding not rated good
Reasoning & maths not rated usable
Agentic work not rated good

Run Ling 3.0 flash on the Framework Desktop, 128GB · or GLM-4.5-Air

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

Side by side on the Framework Desktop, 128GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Strix Halo Framework Desktop, 128GB, at $3,449.

Ling 3.0 flashGLM-4.5-Air
Speed at 32k20 tok/s estimated10 tok/s estimated
Pay-back on this machinePays back in 4,468 yearsPays back in 139 years
API cost per month$0.36$2.66

Run Ling 3.0 flash on the Framework Desktop, 128GB · or GLM-4.5-Air

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Framework Desktop, 128GB. GLM-4.5-Air pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.

A day's useLing 3.0 flashGLM-4.5-Air
50ka few chats a day44,678 years1,392 years
200klight assistant use11,170 years348 years
1Ma moderate coding-assistant day2,234 years70 years
4Mheavy coding with an agent558 years17 years
20Magents running most of the day112 years5.0 yearsits ceiling

The Framework Desktop, 128GB generates at most 13.8M tokens a day on GLM-4.5-Air, so that column's figure at 20M tokens a day is for the most it can do, not for the whole of what was asked.

Run Ling 3.0 flash at 20M tokens a day · or GLM-4.5-Air

Memory is not what separates them

Ling 3.0 flash needs 82 GB of memory at 32k of context and GLM-4.5-Air needs 79 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

Ling 3.0 flash is also head to head with Tencent Hy3 above it on the leaderboard and MiniMax M2.7 below it. One more model needs much the same memory: Devstral 2 123B. GLM-4.5-Air is also head to head with Granite 4.2 8B above it on the leaderboard and Mistral Small 4 (119B-2603) below it. One more model needs much the same memory: Qwen3.5 122B-A10B.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Ling 3.0 flash · every machine that runs GLM-4.5-Air · every other match-up · both against the frontier · the quickest pay-back at each level of use