Sunk Cost sunkcost.ai Data checked 2026-09-03

What can you run with 512 GB of memory?

A model does not get the 512 GB. It gets 384 GB on both machines here sold with 512 GB, because the system keeps the rest. That holds every one of the 39 current models at 32k of context, and the strongest of them is GLM-5.3-Flash.

What a model gets384 GB of the 512 GB. Each machine's page says where the rest goes.
Models that fit at 32kevery one of the 39 current ones, at the quantisation this site lists each at.
The strongest of themGLM-5.3-Flash, sonnet-class.
Cheapest machine at this sizeMac Studio M3 Ultra, 512GB at $9,499, which is what it launched at.

Run the numbers on GLM-5.3-Flash at 512 GB

Jump to: The machines sold with 512 GB · Models that fit in 512 GB · What 512 GB adds over 256 GB · How much context 512 GB leaves room for · Does a 512 GB machine pay for itself?

The machines sold with 512 GB

There are two machines here with 512 GB, and each hands a model the same amount. The fourth column is what each one holds at 32k of context, and it is the number that machine's own page prints.

MachineWhat a model getsPriceModels it holds at 32kStrongest of them
Mac Studio M3 Ultra, 512GBprevious 384 GB $9,499 39 GLM-5.3-Flash Run the numbers
Mac Studio M5 Ultra, 512GB 384 GB not published 39 GLM-5.3-Flash Run the numbers

List prices, and the memory each maker publishes. A machine marked previous is one that is no longer sold, priced at what it launched at.

Models that fit in 512 GB

Every current model a 512 GB machine here holds at 32k of context, strongest first. The third column is the whole job at once: the weights plus the cache for 32k of context. The fourth says how many of the machines at this size have room for it, which is the part a size on its own cannot tell you. Speeds are on the Mac Studio M5 Ultra, 512GB, the machine this list is cut from; speed follows memory bandwidth rather than memory size, so another machine at 512 GB runs the same model at its own rate, and its page gives every one of them.

ModelParametersNeeds at 32kMachines at 512 GBSpeed on the Mac Studio M5 Ultra, 512GB
GLM-5.3-FlashUD-Q4_K_M 321.3B 189.5 GB 2 of 2 32 tok/s estimated Run the numbers
Qwen3.8 Flash NextQ4_K_M 180B 120.5 GB 2 of 2 74 tok/s estimated Run the numbers
DeepSeek V4-FlashUD-Q4_K_M 284B 155.3 GB 2 of 2 49 tok/s estimated Run the numbers
Qwen3.8 27BQ4_K_M 27.8B 18.6 GB 2 of 2 48 tok/s estimated Run the numbers
Tencent Hy3Q4_K_M 295B 192.9 GB 2 of 2 15 tok/s estimated Run the numbers
Inkling SmallUD-Q4_K_M 266B 163.6 GB 2 of 2 43 tok/s estimated Run the numbers
Ling 3.0 flashQ4_K_M 124B 81.6 GB 2 of 2 52 tok/s estimated Run the numbers
MiniMax M2.7Q4_K_M 228.7B 148.5 GB 2 of 2 25 tok/s estimated Run the numbers
Qwen3.6 27BQ4_K_M 27.8B 19 GB 2 of 2 47 tok/s estimated Run the numbers
Qwen3.6 35B-A3BQ4_K_M 36B 22.8 GB 2 of 2 143 tok/s estimated Run the numbers
Muse Glimmer 30BQ4_K_M 29.8B 17.8 GB 2 of 2 50 tok/s estimated Run the numbers
Gemma 4 26B-A4BQ4_K_M 25.2B 17.8 GB 2 of 2 105 tok/s estimated Run the numbers
Qwen3.5 122B-A10BUD-Q4_K_M 125.1B 79.1 GB 2 of 2 51 tok/s estimated Run the numbers
Granite 4.2 30BQ4_K_M 29.3B 26.3 GB 2 of 2 34 tok/s estimated Run the numbers
Gemma 4 31B itQ4_K_M 31.3B 25.8 GB 2 of 2 35 tok/s estimated Run the numbers
GLM-4.7-FlashQ4_K_M 31.2B 20.1 GB 2 of 2 102 tok/s estimated Run the numbers
Nemotron 3.5 Lightning 30B-A3BQ4_K_M 31.6B 25.7 GB 2 of 2 137 tok/s estimated Run the numbers
Gemma 4 12BQ4_K_M 12B 8 GB 2 of 2 113 tok/s estimated Run the numbers
Qwen3.5 9BQ4_K_M 9.7B 6.8 GB 2 of 2 133 tok/s estimated Run the numbers
Qwen3.5 4BQ4_K_M 4.7B 3.8 GB 2 of 2 236 tok/s estimated Run the numbers
MiniCPM5 2BQ4_K_M 2.5B 3 GB 2 of 2 303 tok/s estimated Run the numbers
gpt-oss-120bMXFP4 116.8B 64.6 GB 2 of 2 90 tok/s estimated Run the numbers
Granite 4.2 8BQ4_K_M 8B 10.7 GB 2 of 2 84 tok/s estimated Run the numbers
Ling 3.0 tinyQ4_K_M 7.9B 6.4 GB 2 of 2 150 tok/s estimated Run the numbers
Mistral Small 4 (119B-2603)Q4_K_M 119.4B 74.5 GB 2 of 2 81 tok/s estimated Run the numbers
Qwen3-Coder NextQ4_K_M 79.7B 49.2 GB 2 of 2 137 tok/s estimated Run the numbers
Qwen3-Coder 30B-A3BQ4_K_M 30.5B 21.8 GB 2 of 2 69 tok/s estimated Run the numbers
Devstral 2 123BQ4_K_M 125B 86.7 GB 2 of 2 10 tok/s estimated Run the numbers
gpt-oss-20bMXFP4 20.9B 12.9 GB 2 of 2 124 tok/s estimated Run the numbers
Gemma 4 E4BQAT Q4_0 8B 5.7 GB 2 of 2 104 tok/s estimated Run the numbers
Devstral Small 2 24BQ4_K_M 24B 19.7 GB 2 of 2 46 tok/s estimated Run the numbers
LFM2.5 2.6BQ4_K_M 2.7B 2.2 GB 2 of 2 408 tok/s estimated Run the numbers
Ministral 3 14BQ4_K_M 14B 13.6 GB 2 of 2 66 tok/s estimated Run the numbers
Ministral 3 8BQ4_K_M 8.9B 9.8 GB 2 of 2 92 tok/s estimated Run the numbers
Ornith 1.5 35B-A3BQ4_K_M 36B 22.4 GB 2 of 2 145 tok/s estimated Run the numbers
KAT-Coder V2.5 Dev 35B-A3BQ4_K_M 34.7B 22.1 GB 2 of 2 143 tok/s estimated Run the numbers
Laguna XS 2.1Q4_K_M 33.4B 21.7 GB 2 of 2 112 tok/s estimated Run the numbers
Ornith 1.5 9BQ4_K_M 9.7B 6.9 GB 2 of 2 131 tok/s estimated Run the numbers
Spark-X2.5 4BQ4_K_M 4.1B 3.9 GB 2 of 2 233 tok/s estimated Run the numbers

For scale, the calculator starts from 80 tok/s for a hosted API and times a local machine against it. 22 of the 39 above reach it, and the slowest is 10 tok/s.

Superseded models are left out of the count, because what a machine is worth buying for is what you would run on it today; the leaderboard ranks every model this site lists, older ones included. Where the cache figure comes from is one section of the memory guide.

What 512 GB adds over 256 GB

The machine here with 512 GB that hands a model the most of it is the Mac Studio M5 Ultra, 512GB, at 384 GB, and it holds 39 of the 39. At 256 GB the most is 192 GB, on the Mac Studio M5 Ultra, 256GB, holding 38. The step adds Tencent Hy3. The strongest either way is GLM-5.3-Flash.

The cheapest machine at 512 GB is the Mac Studio M3 Ultra, 512GB at $9,499. The cheapest at 256 GB is the Mac Studio M5 Ultra, 256GB at $10,799, so the larger size is the cheaper of the two here.

Both sides are read off the machine at each size that hands a model the most, so the comparison is the best case against the best case. The table above has every machine at 512 GB and what each of them holds.

How much context 512 GB leaves room for

The cache grows with the window you ask for, so the same machine holds fewer models the longer the context. Every machine at this size hands a model 384 GB, so one column covers all of them. A model is counted only where its own context ceiling reaches that far.

Context384 GB to a modelMac Studio M5 Ultra, 512GB
4k 39
8k 39
16k 39
32k default 39
64k 39
128k 39

Counted over the 39 current models, at the quantisation each is listed at and with the cache at 16 bits. The calculator has a switch for a smaller cache, which changes both what fits and how fast a long window runs.

Does a 512 GB machine pay for itself?

On the Mac Studio M3 Ultra, 512GB at $9,499, the cheapest machine at 512 GB with a published price, the model that pays it back soonest at 500k tokens a day is Inkling Small, in 111 years. At 20M tokens a day, agents running most of the day, it is Inkling Small in 2.8 years. The Mac Studio M3 Ultra, 512GB is the previous generation, so that price is what it launched at.

The sum is the same everywhere on this site: what the same work costs to rent, less what the electricity costs to generate it, against the price of the machine. The best buys rank the quickest pay-back at every level of use, and what it costs a month puts the machine and the API bill in the same shape.

Put your own usage in

Every figure is at 32k of context unless the row says otherwise, with the cache at 16 bits and each model at the quantisation this site lists it at. Weights are the published file sizes on each model's page, and the cache is worked out from the architecture recorded there. What a model gets is the memory the GPU can address, which each machine's page explains. Speeds say whether anybody measured them; where they were not, they are worked out from memory bandwidth. To change the context, the quantisation or the price you would pay, open the calculator. For the sizes either side of this one, the memory guide has the ladder in full.