Sunk Cost sunkcost.ai Data checked 2026-09-03

What can you run with 16 GB of memory?

A model does not get the 16 GB. On the Mac mini M6, 16GB it gets 10.5 GB and on the GeForce RTX 4080, 16GB it gets 15 GB, because the system keeps a share of one and the card keeps a margin free on the other. That is 10 of the 39 current models at 32k of context on the first and 13 on the second, and the strongest of them is Gemma 4 12B.

What a model gets10.5 GB to 15 GB of the 16 GB, depending on the machine. Each machine's page says where the rest goes.
Models that fit at 32k10 to 13 of the 39 current ones, at the quantisation this site lists each at.
The strongest of themGemma 4 12B, below every hosted tier.
Cheapest machine at this sizeMac mini M6, 16GB at $899.

Run the numbers on Gemma 4 12B at 16 GB

Jump to: The machines sold with 16 GB · Models that fit in 16 GB · What 24 GB adds over 16 GB · How much context 16 GB leaves room for · Does a 16 GB machine pay for itself?

The machines sold with 16 GB

There are six machines here with 16 GB, and what they hand a model is not the same figure. The fourth column is what each one holds at 32k of context, and it is the number that machine's own page prints.

MachineWhat a model getsPriceModels it holds at 32kStrongest of them
GeForce RTX 4080, 16GBprevious 15 GB $1,199card only 13 Gemma 4 12B Run the numbers
Mac mini M4, 16GBprevious 10.5 GB $599 10 Gemma 4 12B Run the numbers
Mac mini M6, 16GB 10.5 GB $899 10 Gemma 4 12B Run the numbers
MacBook Air M5 (13-inch), 16GB 10.5 GB $1,299 10 Gemma 4 12B Run the numbers
MacBook Air M5 (15-inch), 16GB 10.5 GB $1,499 10 Gemma 4 12B Run the numbers
MacBook Pro M5 (14-inch), 16GB 10.5 GB $1,999 10 Gemma 4 12B Run the numbers

List prices, and the memory each maker publishes. A machine marked previous is one that is no longer sold, priced at what it launched at. The GeForce RTX 4080, 16GB is priced as the card alone, so add the PC around it before comparing it with a complete computer. All seven cards here are ranked by what each one holds.

Models that fit in 16 GB

Every current model the roomiest 16 GB machine here holds at 32k of context, strongest first. The third column is the whole job at once: the weights plus the cache for 32k of context. The fourth says how many of the machines at this size have room for it, which is the part a size on its own cannot tell you. Speeds are on the GeForce RTX 4080, 16GB, the machine this list is cut from, which is no longer sold; speed follows memory bandwidth rather than memory size, so another machine at 16 GB runs the same model at its own rate, and its page gives every one of them.

ModelParametersNeeds at 32kMachines at 16 GBSpeed on the GeForce RTX 4080, 16GB
Gemma 4 12BQ4_K_M 12B 8 GB 6 of 6 74 tok/s estimated Run the numbers
Qwen3.5 9BQ4_K_M 9.7B 6.8 GB 6 of 6 87 tok/s estimated Run the numbers
Qwen3.5 4BQ4_K_M 4.7B 3.8 GB 6 of 6 154 tok/s estimated Run the numbers
MiniCPM5 2BQ4_K_M 2.5B 3 GB 6 of 6 198 tok/s estimated Run the numbers
Granite 4.2 8BQ4_K_M 8B 10.7 GB 1 of 6 55 tok/s estimated Run the numbers
Ling 3.0 tinyQ4_K_M 7.9B 6.4 GB 6 of 6 125 tok/s estimated Run the numbers
gpt-oss-20bMXFP4 20.9B 12.9 GB 1 of 6 103 tok/s measured Run the numbers
Gemma 4 E4BQAT Q4_0 8B 5.7 GB 6 of 6 87 tok/s estimated Run the numbers
LFM2.5 2.6BQ4_K_M 2.7B 2.2 GB 6 of 6 266 tok/s estimated Run the numbers
Ministral 3 14BQ4_K_M 14B 13.6 GB 1 of 6 43 tok/s estimated Run the numbers
Ministral 3 8BQ4_K_M 8.9B 9.8 GB 6 of 6 60 tok/s estimated Run the numbers
Ornith 1.5 9BQ4_K_M 9.7B 6.9 GB 6 of 6 86 tok/s estimated Run the numbers
Spark-X2.5 4BQ4_K_M 4.1B 3.9 GB 6 of 6 152 tok/s estimated Run the numbers

For scale, the calculator starts from 80 tok/s for a hosted API and times a local machine against it. 9 of the 13 above reach it, and the slowest is 43 tok/s.

Superseded models are left out of the count, because what a machine is worth buying for is what you would run on it today; the leaderboard ranks every model this site lists, older ones included. Where the cache figure comes from is one section of the memory guide.

What 24 GB adds over 16 GB

The machine here with 24 GB that hands a model the most of it is the GeForce RTX 3090, 24GB, at 23 GB, and it holds 24 of the 39. At 16 GB the most is 15 GB, on the GeForce RTX 4080, 16GB, holding 13. The step adds Qwen3.8 27B, Qwen3.6 27B, Qwen3.6 35B-A3B and Muse Glimmer 30B, and seven more. The strongest goes from Gemma 4 12B to Qwen3.8 27B.

The cheapest machine at 24 GB is the Mac mini M6, 24GB at $1,099. The cheapest at 16 GB is the Mac mini M6, 16GB at $899, so the step costs $200.

Both sides are read off the machine at each size that hands a model the most, so the comparison is the best case against the best case. The table above has every machine at 16 GB and what each of them holds.

How much context 16 GB leaves room for

The cache grows with the window you ask for, so the same machine holds fewer models the longer the context. One column here for each amount a 16 GB machine hands over. A model is counted only where its own context ceiling reaches that far.

Context10.5 GB to a modelMac mini M6, 16GB15 GB to a modelGeForce RTX 4080, 16GB
4k 1213
8k 1213
16k 1113
32k default 1013
64k 911
128k 89

Counted over the 39 current models, at the quantisation each is listed at and with the cache at 16 bits. The calculator has a switch for a smaller cache, which changes both what fits and how fast a long window runs.

Does a 16 GB machine pay for itself?

On the Mac mini M6, 16GB at $899, the cheapest machine at 16 GB with a published price, the model that pays it back soonest at 500k tokens a day is Ministral 3 8B, in 37 years. At 20M tokens a day, agents running most of the day, it is Ministral 3 8B in 14 months. The Mac mini M6, 16GB cannot generate 20M tokens in a day; it manages 16.2M, so that figure is for the most it can do.

The sum is the same everywhere on this site: what the same work costs to rent, less what the electricity costs to generate it, against the price of the machine. The best buys rank the quickest pay-back at every level of use, and what it costs a month puts the machine and the API bill in the same shape.

Put your own usage in

Every figure is at 32k of context unless the row says otherwise, with the cache at 16 bits and each model at the quantisation this site lists it at. Weights are the published file sizes on each model's page, and the cache is worked out from the architecture recorded there. What a model gets is the memory the GPU can address, which each machine's page explains. Speeds say whether anybody measured them; where they were not, they are worked out from memory bandwidth. To change the context, the quantisation or the price you would pay, open the calculator. For the sizes either side of this one, the memory guide has the ladder in full.