Qwen3 8B vs Qwen3.5 9B
Qwen3.5 9B scores higher on the intelligence index, 14 against 5. Both take the same machine to start: the cheapest here that runs either is the Mac mini M6, 16GB, at $899. On it, Qwen3.5 9B is about 1.4× quicker: 17 tok/s against 12, both estimated from memory bandwidth. At 500k tokens a day the Mac mini M6, 16GB pays for itself in 40 years running Qwen3 8B, against 69 years running Qwen3.5 9B.
| Qwen3 8B | Qwen3.5 9B | |
|---|---|---|
| Intelligence index | 5 | 14 |
| Class | Below every hosted tier | Below every hosted tier |
| Weights | 5.0 GB | 5.7 GB |
| Needs at 32k | 9.9 GB | 6.8 GB |
| Quantisation | Q4_K_M | Q4_K_M |
| Parameters | 8.2B | 9.7B |
| Max context | 40k | 256k |
| API price per 1M | $0.117 in / $0.455 out | $0.08 in / $0.13 out |
| Licence | Apache 2.0 | Apache 2.0 |
| Machines here that run it | 37 of 37 | 37 of 37 |
| Cheapest machine that runs it | Mac mini M6, 16GB $899 | Mac mini M6, 16GB $899 |
| Summarising | good | not rated |
| Translation | usable | not rated |
| Everyday coding | usable | not rated |
| Reasoning & maths | don’t | not rated |
| Agentic work | don’t | not rated |
Run Qwen3 8B on the Mac mini M6, 16GB · or Qwen3.5 9B
Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.
What the newer model changes
Qwen3 8B is last generation. Qwen3.5 9B is the current dense Qwen model nearest it in size, 9.7B against 8.2B. On the intelligence index it scores 14 where Qwen3 8B scores 5. It asks less of the machine: 6.8 GB at 32k of context against 9.9 GB. The weights are 5.7 GB against 5.0 GB, and the cache at that window is 1.1 GB against 4.8 GB. Every one of the 37 machines priced here runs both, from the Mac mini M6, 16GB at $899.
Qwen3.5 9B takes 256k of context where Qwen3 8B stops at 40k, which is a ceiling rather than a setting: what you actually get is whatever the machine has room for.
Side by side on the Mac mini M6, 16GB
The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Mac mini M6, 16GB, at $899.
| Qwen3 8B | Qwen3.5 9B | |
|---|---|---|
| Speed at 32k | 12 tok/s estimated | 17 tok/s estimated |
| Pay-back on this machine | Pays back in 40 years | Pays back in 69 years |
| API cost per month | $2.10 | $1.27 |
Run Qwen3 8B on the Mac mini M6, 16GB · or Qwen3.5 9B
How much use it takes to pay for the machine
Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Mac mini M6, 16GB. Qwen3 8B pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.
| A day's use | Qwen3 8B | Qwen3.5 9B |
|---|---|---|
| 50ka few chats a day | 405 years | 685 years |
| 200klight assistant use | 101 years | 171 years |
| 1Ma moderate coding-assistant day | 20 years | 34 years |
| 4Mheavy coding with an agent | 5.1 years | 8.6 years |
| 20Magents running most of the day | 15 monthsits ceiling | 21 months |
The Mac mini M6, 16GB generates at most 16.1M tokens a day on Qwen3 8B, so that column's figure at 20M tokens a day is for the most it can do, not for the whole of what was asked.
Run Qwen3 8B at 20M tokens a day · or Qwen3.5 9B
Memory is not what separates them
Qwen3 8B needs 9.9 GB of memory at 32k of context and Qwen3.5 9B needs 6.8 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.
Qwen3 8B is also head to head with Ministral 3 14B above it on the leaderboard and Gemma 3 27B it below it. Qwen3.5 9B is also head to head with Gemma 4 12B above it on the leaderboard and Nemotron 3.5 Lightning 30B-A3B below it.
The assumptions behind both columns
Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.
More head to head: every machine that runs Qwen3 8B · every machine that runs Qwen3.5 9B · every other match-up · both against the frontier · the quickest pay-back at each level of use