Sunk Cost sunkcost.ai Data checked 2026-09-03

Which graphics card should you buy for local LLMs?

Memory decides what you can run; bandwidth decides how fast it runs. Seven cards are priced here, from $329 to $18,000 for the card alone, and they are closer together than that gap suggests. The strongest open model any of them holds is Qwen3.8 27B, and the cheapest card that holds it is the Radeon AI PRO R9700, 32GB at $1,299, card only. The RTX PRO 6000 Blackwell, 96GB at $18,000 runs more models and runs them faster. It does not run a better one.

What decides itMemory first: the weights and the cache for your context window have to fit in the card at the same time. Bandwidth second: it sets the speed.
Best model a card runsQwen3.8 27B, Sonnet-class. Five of the seven cards hold it at 32k context, the cheapest being the Radeon AI PRO R9700, 32GB at $1,299, card only.
Best open model there isGLM-5.3-Flash, 189 GB of weights. No card here holds it. The cheapest machine that does is the Mac Studio M5 Ultra, 256GB at $10,799.
Does one pay for itselfNot at ordinary use. At 20M tokens a day — agents running most of the day — the quickest here is the GeForce RTX 3060, 12GB, in 4.0 months on Ministral 3 8B — the cheapest card working a small model, not the best card working a good one.

Run the numbers on the Radeon AI PRO R9700, 32GB with Qwen3.8 27B

Every card here, side by side

What each one holds at 32k context, counted against the 39 current models, with the strongest of them and how fast it runs. Each price opens the calculator on that card and that model.

CardPriceUsable memoryBandwidthModels it runsStrongest of themSpeed on it
RTX PRO 6000 Blackwell, 96GB $18,000card only 95 GB 1,792 GB/s 33 of 39 Qwen3.8 27B Sonnet-class 78 tok/s estimated
Radeon AI PRO R9700, 32GB $1,299card only 31 GB 640 GB/s 27 of 39 Qwen3.8 27B Sonnet-class 30 tok/s estimated
GeForce RTX 5090, 32GB $1,999card only 31 GB 1,792 GB/s 27 of 39 Qwen3.8 27B Sonnet-class 66 tok/s estimated
GeForce RTX 3090, 24GB $1,499card only, at launch 23 GB 936 GB/s 24 of 39 Qwen3.8 27B Sonnet-class 38 tok/s estimated
GeForce RTX 4090, 24GB $1,599card only, at launch 23 GB 1,008 GB/s 24 of 39 Qwen3.8 27B Sonnet-class 44 tok/s estimated
GeForce RTX 4080, 16GB $1,199card only, at launch 15 GB 717 GB/s 13 of 39 Gemma 4 12B Below every hosted tier 74 tok/s estimated
GeForce RTX 3060, 12GB $329card only, at launch 11 GB 360 GB/s 11 of 39 Gemma 4 12B Below every hosted tier 38 tok/s estimated

Prices are the card on its own: none of them includes the computer you put it in. Four of the seven are previous-generation cards priced at what they launched at, which is not what you would pay for one today — each card's own page says what to enter instead. Usable memory is what a model gets after the card's own overhead, and speed is on the strongest model in the row, so the column is not a race between equals.

Memory is the gate

A card runs a model or it does not, and nothing about the card changes that except how much memory it has. Qwen3.8 27B is 16 GB of weights before a single token of context, and the cache on top grows with every token you keep. That sum, and not the price, is what puts a model on a card. How the two add up is a page of its own.

It is also where these cards stop. Six of the 34 current models with an index score fit none of the seven cards, and the three strongest open models on the site are among them: GLM-5.3-Flash, Qwen3.8 Flash Next and DeepSeek V4-Flash. The dearest card on the list, 95 GB of memory for $18,000, holds none of them either. For a model that size you are buying unified memory instead: GLM-5.3-Flash runs on the Mac Studio M5 Ultra, 256GB at $10,799, and on nothing cheaper.

Bandwidth is the speed

Once a model fits, the card reads every active weight for every token it writes, so tokens a second tracks memory bandwidth more closely than anything else on the spec sheet. On Qwen3.8 27B the RTX PRO 6000 Blackwell, 96GB does 78 tok/s estimated at 1,792 GB/s, against 30 tok/s estimated for the Radeon AI PRO R9700, 32GB at 640 GB/s. Same model, same context, 2.8× the bandwidth.

Every speed on this page is estimated rather than measured, but not out of thin air: each card's estimate is calibrated against that card's own measured llama.cpp runs, which are listed with their sources on its page. Estimates for small models run high, so read them as an upper bound rather than a promise.

What the same money buys with a computer around it

A card's price here buys the card. You still need the machine it goes in, and this site does not guess at what you would build, so every figure above leaves that cost out. What it can show is what the same money already holds when the computer comes with it. Each card is set against the complete machine nearest it in price, with both prices printed, so a near miss is visible rather than smoothed over.

CardIts memoryModelsNearest complete computerIts memoryModels
RTX PRO 6000 Blackwell, 96GB $18,000, card only 95 GB 33 Mac Studio M5 Ultra, 256GB $10,799 192 GB 38
Radeon AI PRO R9700, 32GB $1,299, card only 31 GB 27 Mac mini M6, 32GB $1,299 21 GB 19
GeForce RTX 5090, 32GB $1,999, card only 31 GB 27 MacBook Pro M5 (14-inch), 16GB $1,999 10.5 GB 10
GeForce RTX 3090, 24GB $1,499, card only 23 GB 24 MacBook Air M5 (15-inch), 16GB $1,499 10.5 GB 10
GeForce RTX 4090, 24GB $1,599, card only 23 GB 24 Mac mini M5 Pro, 24GB $1,699 16 GB 13
GeForce RTX 4080, 16GB $1,199, card only 15 GB 13 Framework Desktop, 32GB $1,269 24 GB 24
GeForce RTX 3060, 12GB $329, card only 11 GB 11 Mac mini M6, 16GB $899 10.5 GB 10

On two of the seven the complete computer holds more models than the card does: the Mac Studio M5 Ultra, 256GB holds 38 where the RTX PRO 6000 Blackwell, 96GB holds 33, and the Framework Desktop, 32GB holds 24 where the GeForce RTX 4080, 16GB holds 13. The widest gap is the GeForce RTX 4080, 16GB: 11 models fewer than the Framework Desktop, 32GB, which costs $70 more and is a computer rather than a part for one. What the card has instead is bandwidth: 717 GB/s against 256 GB/s.

Per gigabyte, most of these cards are the expensive way round. The GeForce RTX 3060, 12GB is $30 for each usable gigabyte and the RTX PRO 6000 Blackwell, 96GB is $189. The cheapest gigabyte in a complete computer is the Strix Halo Corsair AI Workstation 300, 64GB at $35, and it arrives with the computer attached. One card beats that — the GeForce RTX 3060, 12GB at $30 — and it is the smallest card here. Divide the price by the usable memory yourself: the two figures are in the tables above. What a card gives back is bandwidth, and only at the top of the range: two of the seven read memory faster than any complete computer here, at 1,792 GB/s against 1,200 GB/s for the Mac Studio M5 Ultra, 96GB. The other five are slower than that machine.

How much use it takes for a card to pay for itself

Everything above is at 500k tokens a day. Pay-back moves with how hard you work a machine, so here is each card at the five levels of use the calculator names, on whichever model it holds that gets there soonest. The levels run from a few chats a day to agents running most of the day.

CardSoonest on50k a day200k a day1M a day4M a day20M a day
RTX PRO 6000 Blackwell, 96GB Qwen3.8 27B 2,273 years568 years114 years28 years5.7 years
Radeon AI PRO R9700, 32GB Qwen3.8 27B 167 years42 years8.3 years2.1 years5.0 months
GeForce RTX 5090, 32GB Qwen3.8 27B 254 years64 years13 years3.2 years7.6 months
GeForce RTX 3090, 24GB Qwen3.8 27B 191 years48 years9.6 years2.4 years5.7 months
GeForce RTX 4090, 24GB Qwen3.8 27B 206 years51 years10 years2.6 years6.2 months
GeForce RTX 4080, 16GB Ministral 3 14B 369 years92 years18 years4.6 years11 months
GeForce RTX 3060, 12GB Ministral 3 8B 134 years34 years6.7 years20 months4.0 months

These are the best case for each card and they still run to years at anything short of constant use, because the pay-back is against renting the same model from an API, and the small models a card holds are the cheap ones to rent. The card price is also the whole of the outlay here, so a real build pays back later than the table says. For the same question asked the other way round — which machine pays back soonest at a level of use — see best buys by usage.

Every figure is at 32k context and at the quantisation named on each model's page, with 15:1 input to output, $0.17 per kWh and today's API prices held flat. Models counted are the 39 current ones, the same set every machine page counts. Card prices are list or launch prices for the card alone, read on 2026-09-03, with the source on each card's page. Set two cards against each other in the head-to-heads, rank the models themselves on the leaderboard, or open the calculator and enter what you actually paid.