DeepSeek V4-Flash vs MiniMax M2.7
DeepSeek V4-Flash scores higher on the intelligence index, 35 against 23. Both take the same machine to start: the cheapest here that runs either is the Mac Studio M5 Ultra, 256GB, at $10,799. On it, DeepSeek V4-Flash is about 2× quicker: 49 tok/s against 25, both estimated from memory bandwidth. At 500k tokens a day the Mac Studio M5 Ultra, 256GB pays for itself in 272 years running MiniMax M2.7, against 1,056 years running DeepSeek V4-Flash.
| DeepSeek V4-Flash | MiniMax M2.7 | |
|---|---|---|
| Intelligence index | 35 | 23 |
| Class | Sonnet-class | Haiku-class |
| Weights | 155 GB | 140 GB |
| Needs at 32k | 155 GB | 149 GB |
| Quantisation | UD-Q4_K_M | Q4_K_M |
| Parameters | 284B (13B active) | 228.7B (10B active) |
| Max context | 1024k | 200k |
| API price per 1M | $0.065 in / $0.18 out | $0.21 in / $0.84 out |
| Licence | MIT | other (see repo) |
| Machines here that run it | 1 of 37 | 1 of 37 |
| Cheapest machine that runs it | Mac Studio M5 Ultra, 256GB $10,799 | Mac Studio M5 Ultra, 256GB $10,799 |
| Summarising | not rated | not rated |
| Translation | not rated | not rated |
| Everyday coding | not rated | not rated |
| Reasoning & maths | not rated | not rated |
| Agentic work | not rated | not rated |
Run DeepSeek V4-Flash on the Mac Studio M5 Ultra, 256GB · or MiniMax M2.7
Ratings are coarse on purpose: they say what a model is usable for, not where it places to the decimal.
Side by side on the Mac Studio M5 Ultra, 256GB
The cheapest machine that runs either model is the same one, so this is the pair doing the same work on the same hardware: Mac Studio M5 Ultra, 256GB, at $10,799.
| DeepSeek V4-Flash | MiniMax M2.7 | |
|---|---|---|
| Speed at 32k | 49 tok/s estimated | 25 tok/s estimated |
| Pay-back on this machine | Pays back in 1,056 years | Pays back in 272 years |
| API cost per month | $1.10 | $3.80 |
Run DeepSeek V4-Flash on the Mac Studio M5 Ultra, 256GB · or MiniMax M2.7
How much use it takes to pay for the machine
Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both models at the five levels of use the calculator names, on the Mac Studio M5 Ultra, 256GB. MiniMax M2.7 pays for it sooner at every level of use, so which of them to run does not turn on how hard you work it.
| A day's use | DeepSeek V4-Flash | MiniMax M2.7 |
|---|---|---|
| 50ka few chats a day | 10,564 years | 2,720 years |
| 200klight assistant use | 2,641 years | 680 years |
| 1Ma moderate coding-assistant day | 528 years | 136 years |
| 4Mheavy coding with an agent | 132 years | 34 years |
| 20Magents running most of the day | 26 years | 6.8 years |
Run DeepSeek V4-Flash at 20M tokens a day · or MiniMax M2.7
Memory is not what separates them
DeepSeek V4-Flash needs 155 GB of memory at 32k of context and MiniMax M2.7 needs 149 GB. Every machine priced here that runs one runs the other, at 32k of context. So the choice between them is what each is good at, how fast it runs and what the same work costs on an API, not what you have to buy to hold it.
DeepSeek V4-Flash is also head to head with Qwen3.8 Flash Next above it on the leaderboard and Qwen3.8 27B below it. One more model needs much the same memory: Inkling Small. MiniMax M2.7 is also head to head with Ling 3.0 flash above it on the leaderboard and Qwen3.6 27B below it. One more model needs much the same memory: Qwen3 235B-A22B Instruct 2507.
The assumptions behind both columns
Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Where nobody rents an open model by the token, its API prices are the nearest hosted model's, named beside them. Machines are the 37 here with a published price that are still sold. Change any of it in the calculator.
More head to head: every machine that runs DeepSeek V4-Flash · every machine that runs MiniMax M2.7 · every other match-up · both against the frontier · the quickest pay-back at each level of use