What can you run with 192 GB of memory?
A model does not get the 192 GB. It gets 160 GB on the one machine here sold with 192 GB, because the system keeps the rest. That holds 36 of the 39 current models at 32k of context, and the strongest of them is Qwen3.8 Flash Next.
Run the numbers on Qwen3.8 Flash Next at 192 GB
Jump to: The machines sold with 192 GB · Models that fit in 192 GB · What 256 GB adds over 192 GB · How much context 192 GB leaves room for · Does a 192 GB machine pay for itself?
The machines sold with 192 GB
One machine here comes with 192 GB, and the fourth column is what it holds at 32k of context: the number its own page prints.
| Machine | What a model gets | Price | Models it holds at 32k | Strongest of them | |
|---|---|---|---|---|---|
| Framework Desktop, 192GB | 160 GB | not published | 36 | Qwen3.8 Flash Next | Run the numbers |
List prices, and the memory each maker publishes. A machine marked previous is one that is no longer sold, priced at what it launched at.
Models that fit in 192 GB
Every current model the one 192 GB machine here holds at 32k of context, strongest first. The third column is the whole job at once: the weights plus the cache for 32k of context. The fourth says how many of the machines at this size have room for it, which is the part a size on its own cannot tell you. Speeds are on the Framework Desktop, 192GB, the machine this list is cut from; speed follows memory bandwidth rather than memory size, so another machine at 192 GB runs the same model at its own rate, and its page gives every one of them.
| Model | Parameters | Needs at 32k | Machines at 192 GB | Speed on the Framework Desktop, 192GB | |
|---|---|---|---|---|---|
| Qwen3.8 Flash NextQ4_K_M | 180B | 120.5 GB | 1 of 1 | 31 tok/s estimated | Run the numbers |
| DeepSeek V4-FlashUD-Q4_K_M | 284B | 155.3 GB | 1 of 1 | 20 tok/s estimated | Run the numbers |
| Qwen3.8 27BQ4_K_M | 27.8B | 18.6 GB | 1 of 1 | 11 tok/s estimated | Run the numbers |
| Ling 3.0 flashQ4_K_M | 124B | 81.6 GB | 1 of 1 | 22 tok/s estimated | Run the numbers |
| MiniMax M2.7Q4_K_M | 228.7B | 148.5 GB | 1 of 1 | 10 tok/s estimated | Run the numbers |
| Qwen3.6 27BQ4_K_M | 27.8B | 19 GB | 1 of 1 | 11 tok/s estimated | Run the numbers |
| Qwen3.6 35B-A3BQ4_K_M | 36B | 22.8 GB | 1 of 1 | 60 tok/s estimated | Run the numbers |
| Muse Glimmer 30BQ4_K_M | 29.8B | 17.8 GB | 1 of 1 | 11 tok/s estimated | Run the numbers |
| Gemma 4 26B-A4BQ4_K_M | 25.2B | 17.8 GB | 1 of 1 | 44 tok/s estimated | Run the numbers |
| Qwen3.5 122B-A10BUD-Q4_K_M | 125.1B | 79.1 GB | 1 of 1 | 21 tok/s estimated | Run the numbers |
| Granite 4.2 30BQ4_K_M | 29.3B | 26.3 GB | 1 of 1 | 7.8 tok/s estimated | Run the numbers |
| Gemma 4 31B itQ4_K_M | 31.3B | 25.8 GB | 1 of 1 | 7.9 tok/s estimated | Run the numbers |
| GLM-4.7-FlashQ4_K_M | 31.2B | 20.1 GB | 1 of 1 | 42 tok/s estimated | Run the numbers |
| Nemotron 3.5 Lightning 30B-A3BQ4_K_M | 31.6B | 25.7 GB | 1 of 1 | 57 tok/s estimated | Run the numbers |
| Gemma 4 12BQ4_K_M | 12B | 8 GB | 1 of 1 | 26 tok/s estimated | Run the numbers |
| Qwen3.5 9BQ4_K_M | 9.7B | 6.8 GB | 1 of 1 | 30 tok/s estimated | Run the numbers |
| Qwen3.5 4BQ4_K_M | 4.7B | 3.8 GB | 1 of 1 | 54 tok/s estimated | Run the numbers |
| MiniCPM5 2BQ4_K_M | 2.5B | 3 GB | 1 of 1 | 69 tok/s estimated | Run the numbers |
| gpt-oss-120bMXFP4 | 116.8B | 64.6 GB | 1 of 1 | 38 tok/s estimated | Run the numbers |
| Granite 4.2 8BQ4_K_M | 8B | 10.7 GB | 1 of 1 | 19 tok/s estimated | Run the numbers |
| Ling 3.0 tinyQ4_K_M | 7.9B | 6.4 GB | 1 of 1 | 62 tok/s estimated | Run the numbers |
| Mistral Small 4 (119B-2603)Q4_K_M | 119.4B | 74.5 GB | 1 of 1 | 34 tok/s estimated | Run the numbers |
| Qwen3-Coder NextQ4_K_M | 79.7B | 49.2 GB | 1 of 1 | 57 tok/s estimated | Run the numbers |
| Qwen3-Coder 30B-A3BQ4_K_M | 30.5B | 21.8 GB | 1 of 1 | 29 tok/s estimated | Run the numbers |
| Devstral 2 123BQ4_K_M | 125B | 86.7 GB | 1 of 1 | 2.4 tok/s estimated | Run the numbers |
| gpt-oss-20bMXFP4 | 20.9B | 12.9 GB | 1 of 1 | 52 tok/s estimated | Run the numbers |
| Gemma 4 E4BQAT Q4_0 | 8B | 5.7 GB | 1 of 1 | 43 tok/s estimated | Run the numbers |
| Devstral Small 2 24BQ4_K_M | 24B | 19.7 GB | 1 of 1 | 10 tok/s estimated | Run the numbers |
| LFM2.5 2.6BQ4_K_M | 2.7B | 2.2 GB | 1 of 1 | 93 tok/s estimated | Run the numbers |
| Ministral 3 14BQ4_K_M | 14B | 13.6 GB | 1 of 1 | 15 tok/s estimated | Run the numbers |
| Ministral 3 8BQ4_K_M | 8.9B | 9.8 GB | 1 of 1 | 21 tok/s estimated | Run the numbers |
| Ornith 1.5 35B-A3BQ4_K_M | 36B | 22.4 GB | 1 of 1 | 60 tok/s estimated | Run the numbers |
| KAT-Coder V2.5 Dev 35B-A3BQ4_K_M | 34.7B | 22.1 GB | 1 of 1 | 60 tok/s estimated | Run the numbers |
| Laguna XS 2.1Q4_K_M | 33.4B | 21.7 GB | 1 of 1 | 47 tok/s estimated | Run the numbers |
| Ornith 1.5 9BQ4_K_M | 9.7B | 6.9 GB | 1 of 1 | 30 tok/s estimated | Run the numbers |
| Spark-X2.5 4BQ4_K_M | 4.1B | 3.9 GB | 1 of 1 | 53 tok/s estimated | Run the numbers |
For scale, the calculator starts from 80 tok/s for a hosted API and times a local machine against it. 1 of the 36 above reaches it, and the slowest is 2.4 tok/s.
Superseded models are left out of the count, because what a machine is worth buying for is what you would run on it today; the leaderboard ranks every model this site lists, older ones included. Where the cache figure comes from is one section of the memory guide.
What 256 GB adds over 192 GB
The machine here with 256 GB that hands a model the most of it is the Mac Studio M5 Ultra, 256GB, at 192 GB, and it holds 38 of the 39. At 192 GB the most is 160 GB, on the Framework Desktop, 192GB, holding 36. The step adds GLM-5.3-Flash and Inkling Small. The strongest goes from Qwen3.8 Flash Next to GLM-5.3-Flash.
Both sides are read off the machine at each size that hands a model the most, so the comparison is the best case against the best case. The table above has every machine at 192 GB and what each of them holds.
How much context 192 GB leaves room for
The cache grows with the window you ask for, so the same machine holds fewer models the longer the context. Every machine at this size hands a model 160 GB, so one column covers all of them. A model is counted only where its own context ceiling reaches that far.
| Context | 160 GB to a modelFramework Desktop, 192GB |
|---|---|
| 4k | 36 |
| 8k | 36 |
| 16k | 36 |
| 32k default | 36 |
| 64k | 36 |
| 128k | 35 |
Counted over the 39 current models, at the quantisation each is listed at and with the cache at 16 bits. The calculator has a switch for a smaller cache, which changes both what fits and how fast a long window runs.
Does a 192 GB machine pay for itself?
No machine here with 192 GB has a published price, so nothing on this page can say what one pays back. The calculator takes yours: put what you would pay beside the machine and it prices the rest.
The sum is the same everywhere on this site: what the same work costs to rent, less what the electricity costs to generate it, against the price of the machine. The best buys rank the quickest pay-back at every level of use, and what it costs a month puts the machine and the API bill in the same shape.
Every figure is at 32k of context unless the row says otherwise, with the cache at 16 bits and each model at the quantisation this site lists it at. Weights are the published file sizes on each model's page, and the cache is worked out from the architecture recorded there. What a model gets is the memory the GPU can address, which each machine's page explains. Speeds say whether anybody measured them; where they were not, they are worked out from memory bandwidth. To change the context, the quantisation or the price you would pay, open the calculator. For the sizes either side of this one, the memory guide has the ladder in full.