Can a NVIDIA RTX PRO 6000 Blackwell, 96GB run local LLMs?
Yes — 33 of the 39 open models on this site fit in its 95 GB of usable memory, the strongest being Qwen3.8 27B. Whether that saves you money is a different question, and the answer is usually no.
Price$18,000
Memory96 GB, about 95 GB of it addressable by the GPU at 1792 GB/s
Best model it runsQwen3.8 27B — Sonnet-class, 78 tok/s
Pay-back against the APIPays back in 227 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Qwen3.8 27BQ4_K_M | 78 tok/s | Sonnet-class | 19 GB | |
| Ling 3.0 flashQ4_K_M | 80 tok/s | Haiku-class | 82 GB | |
| Qwen3.6 27BQ4_K_M | 77 tok/s | Haiku-class | 19 GB | |
| Qwen3.6 35B-A3BQ4_K_M | 221 tok/s | Haiku-class | 23 GB | |
| Muse Glimmer 30BQ4_K_M | 81 tok/s | Haiku-class | 18 GB | |
| Gemma 4 26B-A4BQ4_K_M | 162 tok/s | Haiku-class | 18 GB | |
| Qwen3.5 122B-A10BUD-Q4_K_M | 79 tok/s | Haiku-class | 79 GB | |
| Gemma 4 31B itQ4_K_M | 56 tok/s | Haiku-class | 26 GB | |
| Granite 4.2 30BQ4_K_M | 55 tok/s | Haiku-class | 26 GB | |
| GLM-4.7-FlashQ4_K_M | 157 tok/s | Haiku-class | 20 GB | |
| Gemma 4 12BQ4_K_M | 182 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 215 tok/s | Below every hosted tier | 6.8 GB |
21 more fit; the calculator lists them all.
The specifics
- Chip
- RTX PRO 6000 Blackwell — Blackwell · 96 GB GDDR7 ECC
- Memory bandwidth
- 1792 GB/s
- Usable by the GPU
- 95 GB — Bandwidth: 28 Gbps × 512-bit bus ÷ 8 = 1,792 GB/s (Puget Systems; NVIDIA lists 1792 GB/sec). Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
- Power under load
- 600 W (published) — 600 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. One llama.cpp user saw about 390 W while benchmarking gpt-oss-120b on this card.
- Availability
- Newegg price on 2026-09-16; NVIDIA publishes no list price. Not the 300 W Max-Q version.
- Sources
- source 1, source 2, source 3, source 4, source 5, source 6