Sunk Cost sunkcost.ai Data checked 2026-09-03

NVIDIA RTX PRO 6000 Blackwell, 96GB vs NVIDIA GeForce RTX 5090, 32GB for local AI

The RTX PRO 6000 Blackwell, 96GB holds 33 of the 39 open models here, and the GeForce RTX 5090, 32GB holds 27. The GeForce RTX 5090, 32GB costs $16,001 less, both priced as the card alone, without the PC around either. On Qwen3.8 27B, the strongest model both hold, the RTX PRO 6000 Blackwell, 96GB is about 1.2× faster: 78 tok/s against 66, both estimated from memory bandwidth. The GeForce RTX 5090, 32GB pays for itself sooner, in 25 years against 227 years at 500k tokens a day.

NVIDIA RTX PRO 6000 Blackwell, 96GBNVIDIA GeForce RTX 5090, 32GB
Price$18,000card only$1,999card only
Memory96 GB32 GB
Usable by the GPU95 GB31 GB
Memory bandwidth1792 GB/s1792 GB/s
Power under load600 W575 W
Models that fit3327
Best model it runsQwen3.8 27BQwen3.8 27B
Speed on that model78 tok/s estimated66 tok/s estimated
Pay-back on that modelPays back in 227 yearsPays back in 25 years

Run the numbers on the NVIDIA RTX PRO 6000 Blackwell, 96GB · or the NVIDIA GeForce RTX 5090, 32GB

How much use it takes to pay back

Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both machines on Qwen3.8 27B, the strongest model both hold, at the five levels of use the calculator names. The NVIDIA GeForce RTX 5090, 32GB pays back sooner at every level of use, so this is not a choice that turns on how hard you work it.

A day's useNVIDIA RTX PRO 6000 Blackwell, 96GBNVIDIA GeForce RTX 5090, 32GB
50ka few chats a day2,273 years254 years
200klight assistant use568 years64 years
1Ma moderate coding-assistant day114 years13 years
4Mheavy coding with an agent28 years3.2 years
20Magents running most of the day5.7 years7.6 months

Run the NVIDIA RTX PRO 6000 Blackwell, 96GB at 20M tokens a day · or the NVIDIA GeForce RTX 5090, 32GB

What the extra memory buys

The NVIDIA RTX PRO 6000 Blackwell, 96GB holds 6 models the NVIDIA GeForce RTX 5090, 32GB cannot at 32k of context. The strongest of them are what the difference in memory actually buys.

ModelWeightsNeeds at 32kOn the NVIDIA RTX PRO 6000 Blackwell, 96GB
Ling 3.0 flashHaiku-class78 GB82 GB80 tok/s estimated
Qwen3.5 122B-A10BHaiku-class78 GB79 GB79 tok/s estimated
gpt-oss-120bBelow every hosted tier63 GB65 GB136 tok/s measured
Mistral Small 4 (119B-2603)Below every hosted tier74 GB75 GB124 tok/s estimated
Qwen3-Coder NextBelow every hosted tier48 GB49 GB211 tok/s estimated
Devstral 2 123BBelow every hosted tier75 GB87 GB17 tok/s estimated

The assumptions behind both columns

Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Graphics cards are priced as the card alone, so add the PC around one before comparing it with a complete computer. Change any of it in the calculator.

More head to head: everything the NVIDIA RTX PRO 6000 Blackwell, 96GB runs · everything the NVIDIA GeForce RTX 5090, 32GB runs · every other match-up · the quickest pay-back at each level of use · every model against the frontier