What can you run with 96 GB of memory?
A model does not get the 96 GB. On the Mac Studio M5 Ultra, 96GB it gets 72 GB and on the RTX PRO 6000 Blackwell, 96GB it gets 95 GB, because the system keeps a share of one and the card keeps a margin free on the other. That is 29 of the 39 current models at 32k of context on the first and 33 on the second, and the strongest of them is Qwen3.8 27B.
Run the numbers on Qwen3.8 27B at 96 GB
Jump to: The machines sold with 96 GB · Models that fit in 96 GB · What 128 GB adds over 96 GB · How much context 96 GB leaves room for · Does a 96 GB machine pay for itself?
The machines sold with 96 GB
There are three machines here with 96 GB, and what they hand a model is not the same figure. The fourth column is what each one holds at 32k of context, and it is the number that machine's own page prints.
| Machine | What a model gets | Price | Models it holds at 32k | Strongest of them | |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell, 96GB | 95 GB | $18,000card only | 33 | Qwen3.8 27B | Run the numbers |
| Mac Studio M3 Ultra, 96GBprevious | 72 GB | $3,999 | 29 | Qwen3.8 27B | Run the numbers |
| Mac Studio M5 Ultra, 96GB | 72 GB | $5,499 | 29 | Qwen3.8 27B | Run the numbers |
List prices, and the memory each maker publishes. A machine marked previous is one that is no longer sold, priced at what it launched at. The RTX PRO 6000 Blackwell, 96GB is priced as the card alone, so add the PC around it before comparing it with a complete computer. All seven cards here are ranked by what each one holds.
Models that fit in 96 GB
Every current model the roomiest 96 GB machine here holds at 32k of context, strongest first. The third column is the whole job at once: the weights plus the cache for 32k of context. The fourth says how many of the machines at this size have room for it, which is the part a size on its own cannot tell you. Speeds are on the RTX PRO 6000 Blackwell, 96GB, the machine this list is cut from; speed follows memory bandwidth rather than memory size, so another machine at 96 GB runs the same model at its own rate, and its page gives every one of them.
| Model | Parameters | Needs at 32k | Machines at 96 GB | Speed on the RTX PRO 6000 Blackwell, 96GB | |
|---|---|---|---|---|---|
| Qwen3.8 27BQ4_K_M | 27.8B | 18.6 GB | 3 of 3 | 78 tok/s estimated | Run the numbers |
| Ling 3.0 flashQ4_K_M | 124B | 81.6 GB | 1 of 3 | 80 tok/s estimated | Run the numbers |
| Qwen3.6 27BQ4_K_M | 27.8B | 19 GB | 3 of 3 | 77 tok/s estimated | Run the numbers |
| Qwen3.6 35B-A3BQ4_K_M | 36B | 22.8 GB | 3 of 3 | 221 tok/s estimated | Run the numbers |
| Muse Glimmer 30BQ4_K_M | 29.8B | 17.8 GB | 3 of 3 | 81 tok/s estimated | Run the numbers |
| Gemma 4 26B-A4BQ4_K_M | 25.2B | 17.8 GB | 3 of 3 | 162 tok/s estimated | Run the numbers |
| Qwen3.5 122B-A10BUD-Q4_K_M | 125.1B | 79.1 GB | 1 of 3 | 79 tok/s estimated | Run the numbers |
| Granite 4.2 30BQ4_K_M | 29.3B | 26.3 GB | 3 of 3 | 55 tok/s estimated | Run the numbers |
| Gemma 4 31B itQ4_K_M | 31.3B | 25.8 GB | 3 of 3 | 56 tok/s estimated | Run the numbers |
| GLM-4.7-FlashQ4_K_M | 31.2B | 20.1 GB | 3 of 3 | 157 tok/s estimated | Run the numbers |
| Nemotron 3.5 Lightning 30B-A3BQ4_K_M | 31.6B | 25.7 GB | 3 of 3 | 212 tok/s estimated | Run the numbers |
| Gemma 4 12BQ4_K_M | 12B | 8 GB | 3 of 3 | 182 tok/s estimated | Run the numbers |
| Qwen3.5 9BQ4_K_M | 9.7B | 6.8 GB | 3 of 3 | 215 tok/s estimated | Run the numbers |
| Qwen3.5 4BQ4_K_M | 4.7B | 3.8 GB | 3 of 3 | 381 tok/s estimated | Run the numbers |
| MiniCPM5 2BQ4_K_M | 2.5B | 3 GB | 3 of 3 | 489 tok/s estimated | Run the numbers |
| gpt-oss-120bMXFP4 | 116.8B | 64.6 GB | 3 of 3 | 136 tok/s measured | Run the numbers |
| Granite 4.2 8BQ4_K_M | 8B | 10.7 GB | 3 of 3 | 135 tok/s estimated | Run the numbers |
| Ling 3.0 tinyQ4_K_M | 7.9B | 6.4 GB | 3 of 3 | 231 tok/s estimated | Run the numbers |
| Mistral Small 4 (119B-2603)Q4_K_M | 119.4B | 74.5 GB | 1 of 3 | 124 tok/s estimated | Run the numbers |
| Qwen3-Coder NextQ4_K_M | 79.7B | 49.2 GB | 3 of 3 | 211 tok/s estimated | Run the numbers |
| Qwen3-Coder 30B-A3BQ4_K_M | 30.5B | 21.8 GB | 3 of 3 | 106 tok/s estimated | Run the numbers |
| Devstral 2 123BQ4_K_M | 125B | 86.7 GB | 1 of 3 | 17 tok/s estimated | Run the numbers |
| gpt-oss-20bMXFP4 | 20.9B | 12.9 GB | 3 of 3 | 207 tok/s measured | Run the numbers |
| Gemma 4 E4BQAT Q4_0 | 8B | 5.7 GB | 3 of 3 | 161 tok/s estimated | Run the numbers |
| Devstral Small 2 24BQ4_K_M | 24B | 19.7 GB | 3 of 3 | 74 tok/s estimated | Run the numbers |
| LFM2.5 2.6BQ4_K_M | 2.7B | 2.2 GB | 3 of 3 | 658 tok/s estimated | Run the numbers |
| Ministral 3 14BQ4_K_M | 14B | 13.6 GB | 3 of 3 | 107 tok/s estimated | Run the numbers |
| Ministral 3 8BQ4_K_M | 8.9B | 9.8 GB | 3 of 3 | 149 tok/s estimated | Run the numbers |
| Ornith 1.5 35B-A3BQ4_K_M | 36B | 22.4 GB | 3 of 3 | 224 tok/s estimated | Run the numbers |
| KAT-Coder V2.5 Dev 35B-A3BQ4_K_M | 34.7B | 22.1 GB | 3 of 3 | 220 tok/s estimated | Run the numbers |
| Laguna XS 2.1Q4_K_M | 33.4B | 21.7 GB | 3 of 3 | 172 tok/s estimated | Run the numbers |
| Ornith 1.5 9BQ4_K_M | 9.7B | 6.9 GB | 3 of 3 | 212 tok/s estimated | Run the numbers |
| Spark-X2.5 4BQ4_K_M | 4.1B | 3.9 GB | 3 of 3 | 376 tok/s estimated | Run the numbers |
For scale, the calculator starts from 80 tok/s for a hosted API and times a local machine against it. 26 of the 33 above reach it, and the slowest is 17 tok/s.
Superseded models are left out of the count, because what a machine is worth buying for is what you would run on it today; the leaderboard ranks every model this site lists, older ones included. Where the cache figure comes from is one section of the memory guide.
What 128 GB adds over 96 GB
The machine here with 128 GB that hands a model the most of it is the DGX Spark, 128GB, at 119.5 GB, and it holds 33 of the 39. At 96 GB the most is 95 GB, on the RTX PRO 6000 Blackwell, 96GB, holding 33. Nothing new fits. What the step buys is a longer window and room to work, not a model that was out of reach. The strongest either way is Qwen3.8 27B.
The cheapest machine at 128 GB is the Strix Halo Framework Desktop, 128GB at $3,449. The cheapest at 96 GB is the Mac Studio M5 Ultra, 96GB at $5,499, so the larger size is the cheaper of the two here.
Both sides are read off the machine at each size that hands a model the most, so the comparison is the best case against the best case. The table above has every machine at 96 GB and what each of them holds.
How much context 96 GB leaves room for
The cache grows with the window you ask for, so the same machine holds fewer models the longer the context. One column here for each amount a 96 GB machine hands over. A model is counted only where its own context ceiling reaches that far.
| Context | 72 GB to a modelMac Studio M5 Ultra, 96GB | 95 GB to a modelRTX PRO 6000 Blackwell, 96GB |
|---|---|---|
| 4k | 29 | 33 |
| 8k | 29 | 33 |
| 16k | 29 | 33 |
| 32k default | 29 | 33 |
| 64k | 29 | 32 |
| 128k | 29 | 32 |
Counted over the 39 current models, at the quantisation each is listed at and with the cache at 16 bits. The calculator has a switch for a smaller cache, which changes both what fits and how fast a long window runs.
Does a 96 GB machine pay for itself?
On the Mac Studio M5 Ultra, 96GB at $5,499, the cheapest machine at 96 GB with a published price, the model that pays it back soonest at 500k tokens a day is Qwen3.8 27B, in 68 years. At 20M tokens a day, agents running most of the day, it is Qwen3.8 27B in 21 months.
The sum is the same everywhere on this site: what the same work costs to rent, less what the electricity costs to generate it, against the price of the machine. The best buys rank the quickest pay-back at every level of use, and what it costs a month puts the machine and the API bill in the same shape.
Every figure is at 32k of context unless the row says otherwise, with the cache at 16 bits and each model at the quantisation this site lists it at. Weights are the published file sizes on each model's page, and the cache is worked out from the architecture recorded there. What a model gets is the memory the GPU can address, which each machine's page explains. Speeds say whether anybody measured them; where they were not, they are worked out from memory bandwidth. To change the context, the quantisation or the price you would pay, open the calculator. For the sizes either side of this one, the memory guide has the ladder in full.