Sunk Cost sunkcost.ai Data checked 2026-09-03

Qwen3.5 9B vs Ling 3.0 tiny

Qwen3.5 9B scores higher on the intelligence index, 14 against 12. Both take the same machine to start: the cheapest here that runs either is the Mac mini M6, 16GB, at $899. On it, Ling 3.0 tiny is about 1.1× quicker: 19 tok/s against 17, both estimated from memory bandwidth. At 500k tokens a day the Mac mini M6, 16GB pays for itself in 69 years running Qwen3.5 9B, against 80 years running Ling 3.0 tiny.

Qwen3.5 9BLing 3.0 tiny
Intelligence index1412
ClassBelow every hosted tierBelow every hosted tier
Weights5.7 GB4.8 GB
Needs at 32k6.8 GB6.4 GB
QuantisationQ4_K_MQ4_K_M
Parameters9.7B7.9B (1.3B active)
Max context256k128k
API price per 1M$0.08 in / $0.13 out$0.06 in / $0.25 outpriced as Granite 4.2 8B
LicenceApache 2.0MIT
Machines here that run it37 of 3737 of 37
Cheapest machine that runs itMac mini M6, 16GB $899Mac mini M6, 16GB $899
Summarising not rated not rated
Translation not rated not rated
Everyday coding not rated not rated
Reasoning & maths not rated not rated
Agentic work not rated not rated

Run Qwen3.5 9B on the Mac mini M6, 16GB · or Ling 3.0 tiny

Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.

Side by side on the Mac mini M6, 16GB

The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Mac mini M6, 16GB, at $899.

Qwen3.5 9BLing 3.0 tiny
Speed at 32k17 tok/s estimated19 tok/s estimated
Pay-back on this machinePays back in 69 yearsPays back in 80 years
API cost per month$1.27$1.09priced as Granite 4.2 8B

Run Qwen3.5 9B on the Mac mini M6, 16GB · or Ling 3.0 tiny

How much use it takes to pay for the machine

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Mac mini M6, 16GB. Qwen3.5 9B pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.

A day's useQwen3.5 9BLing 3.0 tiny
50ka few chats a day685 years796 years
200klight assistant use171 years199 years
1Ma moderate coding-assistant day34 years40 years
4Mheavy coding with an agent8.6 years10.0 years
20Magents running most of the day21 months24 months

Run Qwen3.5 9B at 20M tokens a day · or Ling 3.0 tiny

Memory is not what separates them

Qwen3.5 9B needs 6.8 GB of memory at 32k of context and Ling 3.0 tiny needs 6.4 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.

Qwen3.5 9B is also head to head with Gemma 4 12B above it on the leaderboard, Nemotron 3.5 Lightning 30B-A3B below it and Qwen3 8B, the last-generation Qwen nearest it in size. Ling 3.0 tiny is also head to head with gpt-oss-120b above it on the leaderboard and Granite 4.2 8B below it.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.

More head to head: every machine that runs Qwen3.5 9B · every machine that runs Ling 3.0 tiny · every other match-up · both against the frontier · the quickest pay-back at each level of use