Qwen3 235B-A22B Instruct 2507 vs Qwen3.8 Flash Next
Qwen3.8 Flash Next scores higher on the intelligence index, 40 against 13. Both take the same machine to start: the cheapest here that runs either is the Mac Studio M5 Ultra, 256GB, at $10,799. On it, Qwen3.8 Flash Next is about 4.1× quicker: 74 tok/s against 18, both estimated from memory bandwidth. At 500k tokens a day the Mac Studio M5 Ultra, 256GB pays for itself in 372 years running Qwen3.8 Flash Next, against 977 years running Qwen3 235B-A22B Instruct 2507.
| Qwen3 235B-A22B Instruct 2507 | Qwen3.8 Flash Next | |
|---|---|---|
| Intelligence index | 13 | 40 |
| Class | Below every hosted tier | Sonnet-class |
| Weights | 142 GB | 120 GB |
| Needs at 32k | 148 GB | 121 GB |
| Quantisation | Q4_K_M | Q4_K_M |
| Parameters | 235.1B (22B active) | 180B (6B active) |
| Max context | 256k | 256k |
| API price per 1M | $0.0875 in / $0.35 out | $0.15 in / $0.47 out |
| Licence | Apache 2.0 | qwen-community-1.0 |
| Machines here that run it | 1 of 37 | 1 of 37 |
| Cheapest machine that runs it | Mac Studio M5 Ultra, 256GB $10,799 | Mac Studio M5 Ultra, 256GB $10,799 |
| Summarising | good | not rated |
| Translation | good | not rated |
| Everyday coding | good | not rated |
| Reasoning & maths | good | not rated |
| Agentic work | usable | not rated |
Run Qwen3 235B-A22B Instruct 2507 on the Mac Studio M5 Ultra, 256GB · or Qwen3.8 Flash Next
Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.
What the newer model changes
Qwen3 235B-A22B Instruct 2507 is last generation. Qwen3.8 Flash Next is the current Qwen mixture of experts nearest it in size, 180B against 235.1B. On the intelligence index it scores 40 where Qwen3 235B-A22B Instruct 2507 scores 13. It asks less of the machine: 121 GB at 32k of context against 148 GB. The weights are 120 GB against 142 GB, and the cache at that window is 0.9 GB against 6.3 GB. One of the 37 machines priced here runs either of them, and it is the same one: the Mac Studio M5 Ultra, 256GB at $10,799.
Side by side on the Mac Studio M5 Ultra, 256GB
The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Mac Studio M5 Ultra, 256GB, at $10,799.
| Qwen3 235B-A22B Instruct 2507 | Qwen3.8 Flash Next | |
|---|---|---|
| Speed at 32k | 18 tok/s estimated | 74 tok/s estimated |
| Pay-back on this machine | Pays back in 977 years | Pays back in 372 years |
| API cost per month | $1.58 | $2.59 |
Run Qwen3 235B-A22B Instruct 2507 on the Mac Studio M5 Ultra, 256GB · or Qwen3.8 Flash Next
How much use it takes to pay for the machine
Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Mac Studio M5 Ultra, 256GB. Qwen3.8 Flash Next pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.
| A day's use | Qwen3 235B-A22B Instruct 2507 | Qwen3.8 Flash Next |
|---|---|---|
| 50ka few chats a day | 9,774 years | 3,715 years |
| 200klight assistant use | 2,444 years | 929 years |
| 1Ma moderate coding-assistant day | 489 years | 186 years |
| 4Mheavy coding with an agent | 122 years | 46 years |
| 20Magents running most of the day | 24 years | 9.3 years |
Run Qwen3 235B-A22B Instruct 2507 at 20M tokens a day · or Qwen3.8 Flash Next
Memory is not what separates them
Qwen3 235B-A22B Instruct 2507 needs 148 GB of memory at 32k of context and Qwen3.8 Flash Next needs 121 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.
Qwen3 235B-A22B Instruct 2507 is also head to head with Nemotron 3.5 Lightning 30B-A3B above it on the leaderboard and MiniCPM5 2B below it. Qwen3.8 Flash Next is also head to head with GLM-5.3-Flash above it on the leaderboard and DeepSeek V4-Flash below it.
The assumptions behind both columns
Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.
More head to head: every machine that runs Qwen3 235B-A22B Instruct 2507 · every machine that runs Qwen3.8 Flash Next · every other match-up · both against the frontier · the quickest pay-back at each level of use