NVIDIA RTX PRO 6000 Blackwell, 96GB vs NVIDIA GeForce RTX 4090, 24GB for local AI
The RTX PRO 6000 Blackwell, 96GB holds 33 of the 39 open models here, and the GeForce RTX 4090, 24GB holds 24. The GeForce RTX 4090, 24GB costs $16,401 less, both priced as the card alone, without the PC around either. On Qwen3.8 27B, the strongest model both hold, the RTX PRO 6000 Blackwell, 96GB is about 1.8× faster: 78 tok/s against 44, both estimated from memory bandwidth. The GeForce RTX 4090, 24GB pays for itself sooner, in 21 years against 227 years at 500k tokens a day. The GeForce RTX 4090, 24GB is the previous generation, so every figure here for it is priced at what it launched at rather than at a price you can pay today.
| NVIDIA RTX PRO 6000 Blackwell, 96GB | NVIDIA GeForce RTX 4090, 24GB | |
|---|---|---|
| Price | $18,000card only | $1,599card only |
| Memory | 96 GB | 24 GB |
| Usable by the GPU | 95 GB | 23 GB |
| Memory bandwidth | 1792 GB/s | 1008 GB/s |
| Power under load | 600 W | 450 W |
| Models that fit | 33 | 24 |
| Best model it runs | Qwen3.8 27B | Qwen3.8 27B |
| Speed on that model | 78 tok/s estimated | 44 tok/s estimated |
| Pay-back on that model | Pays back in 227 years | Pays back in 21 years |
Run the numbers on the NVIDIA RTX PRO 6000 Blackwell, 96GB · or the NVIDIA GeForce RTX 4090, 24GB
How much use it takes to pay back
Everything above is at 500k tokens a day. Pay-back moves with how much you actually run, so here are both machines on Qwen3.8 27B, the strongest model both hold, at the five levels of use the calculator names. The NVIDIA GeForce RTX 4090, 24GB pays back sooner at every level of use, so this is not a choice that turns on how hard you work it.
| A day's use | NVIDIA RTX PRO 6000 Blackwell, 96GB | NVIDIA GeForce RTX 4090, 24GB |
|---|---|---|
| 50ka few chats a day | 2,273 years | 206 years |
| 200klight assistant use | 568 years | 51 years |
| 1Ma moderate coding-assistant day | 114 years | 10 years |
| 4Mheavy coding with an agent | 28 years | 2.6 years |
| 20Magents running most of the day | 5.7 years | 6.2 months |
Run the NVIDIA RTX PRO 6000 Blackwell, 96GB at 20M tokens a day · or the NVIDIA GeForce RTX 4090, 24GB
What the extra memory buys
The NVIDIA RTX PRO 6000 Blackwell, 96GB holds 9 models the NVIDIA GeForce RTX 4090, 24GB cannot at 32k of context. The strongest of them are what the difference in memory actually buys.
| Model | Weights | Needs at 32k | On the NVIDIA RTX PRO 6000 Blackwell, 96GB |
|---|---|---|---|
| Ling 3.0 flashHaiku-class | 78 GB | 82 GB | 80 tok/s estimated |
| Qwen3.5 122B-A10BHaiku-class | 78 GB | 79 GB | 79 tok/s estimated |
| Gemma 4 31B itHaiku-class | 20 GB | 26 GB | 56 tok/s estimated |
| Granite 4.2 30BHaiku-class | 18 GB | 26 GB | 55 tok/s estimated |
| Nemotron 3.5 Lightning 30B-A3BBelow every hosted tier | 25 GB | 26 GB | 212 tok/s estimated |
| gpt-oss-120bBelow every hosted tier | 63 GB | 65 GB | 136 tok/s measured |
3 more, on the NVIDIA RTX PRO 6000 Blackwell, 96GB page.
The assumptions behind both columns
Both columns use the same usage: 500k tokens a day at 15:1 input to output, 32k of context, $0.17 per kWh, and today's API prices held flat. Speeds marked estimated are worked out from memory bandwidth rather than measured, and pay-back scales with them. Graphics cards are priced as the card alone, so add the PC around one before comparing it with a complete computer. Change any of it in the calculator.
More head to head: everything the NVIDIA RTX PRO 6000 Blackwell, 96GB runs · everything the NVIDIA GeForce RTX 4090, 24GB runs · every other match-up · the quickest pay-back at each level of use · every model against the frontier