Sunk Cost sunkcost.ai Data checked 2026-09-03

Mac Studio M5 Ultra, 96GB vs Strix Halo HP Z2 Mini G1a, 128GB for local AI

The HP Z2 Mini G1a, 128GB holds 33 of the 39 open models here, and the Mac Studio M5 Ultra, 96GB holds 29. The Mac Studio M5 Ultra, 96GB costs $45 less. On Qwen3.8 27B, the strongest model both hold, the Mac Studio M5 Ultra, 96GB is about 4.8× faster: 48 tok/s against 10, both estimated from memory bandwidth. The Mac Studio M5 Ultra, 96GB pays for itself sooner, in 68 years against 73 years at 500k tokens a day.

Mac Studio M5 Ultra, 96GBStrix Halo HP Z2 Mini G1a, 128GB
Price$5,499$5,544
Memory96 GB128 GB
Usable by the GPU72 GB96 GB
Memory bandwidth1200 GB/s256 GB/s
Power under load270 Wstand-in133 Wstand-in
Models that fit2933
Best model it runsQwen3.8 27BQwen3.8 27B
Speed on that model48 tok/s estimated10 tok/s estimated
Pay-back on that modelPays back in 68 yearsPays back in 73 years

Neither power figure above is measured on the machine beside it. Both are stand-ins borrowed from the nearest hardware the data does have, so the gap between them is not a difference between these two machines. Each machine's page names the figure it borrows and why. Every pay-back figure on this page prices its electricity from these numbers.

Run the numbers on the Mac Studio M5 Ultra, 96GB · or the Strix Halo HP Z2 Mini G1a, 128GB

The same money, two different machines

These two cost within $45 of each other: $5,499 for the Mac Studio M5 Ultra, 96GB and $5,544 for the HP Z2 Mini G1a, 128GB. Apple makes one and HP the other. Every other head-to-head on this site holds a piece of the hardware equal and asks what the price gap buys. This one holds the price, so the row that usually carries the answer is the row the two machines agree on, and everything under it is what the same money buys twice.

How much use it takes to pay back

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both machines on Qwen3.8 27B, the strongest model both hold, at the five levels of use the calculator names. The Mac Studio M5 Ultra, 96GB pays back sooner at every level of use, so this is not a choice that turns on how hard you work it.

A day's useMac Studio M5 Ultra, 96GBStrix Halo HP Z2 Mini G1a, 128GB
50ka few chats a day685 years726 years
200klight assistant use171 years181 years
1Ma moderate coding-assistant day34 years36 years
4Mheavy coding with an agent8.6 years9.1 years
20Magents running most of the day21 months2.5 yearsits ceiling

On Qwen3.8 27B the Strix Halo HP Z2 Mini G1a, 128GB generates at most 14.3M tokens a day, so its figure at 20M tokens a day is for the most it can do, not for the whole of what was asked.

Run the Mac Studio M5 Ultra, 96GB at 20M tokens a day · or the Strix Halo HP Z2 Mini G1a, 128GB

What the extra memory buys

The Strix Halo HP Z2 Mini G1a, 128GB holds 4 models the Mac Studio M5 Ultra, 96GB cannot at 32k of context. The strongest of them are what the difference in memory actually buys.

ModelWeightsNeeds at 32kOn the Strix Halo HP Z2 Mini G1a, 128GB
Ling 3.0 flashHaiku-class78 GB82 GB20 tok/s estimated
Qwen3.5 122B-A10BHaiku-class78 GB79 GB20 tok/s estimated
Mistral Small 4 (119B-2603)Below every hosted tier74 GB75 GB32 tok/s estimated
Devstral 2 123BBelow every hosted tier75 GB87 GB2.2 tok/s estimated

The Mac Studio M5 Ultra, 96GB is also head to head with another computer: MacBook Pro M5 Max (16-inch), 64GB. With the same box and the bigger chip: 256GB. With the machine it replaced: Mac Studio M3 Ultra, 96GB.

The HP Z2 Mini G1a, 128GB is also head to head with another computer: Framework Desktop, 128GB.

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Change any of it in the calculator.

More head to head: everything the Mac Studio M5 Ultra, 96GB runs · everything the Strix Halo HP Z2 Mini G1a, 128GB runs · every other match-up · the quickest pay-back at each level of use · every model against the frontier