Which graphics card should you buy for local LLMs?
Memory decides what you can run; bandwidth decides how fast it runs. Seven cards are priced here, from $329 to $18,000 for the card alone, and they are closer together than that gap suggests. The strongest open model any of them holds is Qwen3.8 27B, and the cheapest card that holds it is the Radeon AI PRO R9700, 32GB at $1,299, card only. The RTX PRO 6000 Blackwell, 96GB at $18,000 runs more models and runs them faster. It does not run a better one.
Run the numbers on the Radeon AI PRO R9700, 32GB with Qwen3.8 27B
Every card here, side by side
What each one holds at 32k context, counted against the 39 current models, with the strongest of them and how fast it runs. Each price opens the calculator on that card and that model.
| Card | Price | Usable memory | Bandwidth | Models it runs | Strongest of them | Speed on it |
|---|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell, 96GB | $18,000card only | 95 GB | 1,792 GB/s | 33 of 39 | Qwen3.8 27B Sonnet-class | 78 tok/s estimated |
| Radeon AI PRO R9700, 32GB | $1,299card only | 31 GB | 640 GB/s | 27 of 39 | Qwen3.8 27B Sonnet-class | 30 tok/s estimated |
| GeForce RTX 5090, 32GB | $1,999card only | 31 GB | 1,792 GB/s | 27 of 39 | Qwen3.8 27B Sonnet-class | 66 tok/s estimated |
| GeForce RTX 3090, 24GB | $1,499card only, at launch | 23 GB | 936 GB/s | 24 of 39 | Qwen3.8 27B Sonnet-class | 38 tok/s estimated |
| GeForce RTX 4090, 24GB | $1,599card only, at launch | 23 GB | 1,008 GB/s | 24 of 39 | Qwen3.8 27B Sonnet-class | 44 tok/s estimated |
| GeForce RTX 4080, 16GB | $1,199card only, at launch | 15 GB | 717 GB/s | 13 of 39 | Gemma 4 12B Below every hosted tier | 74 tok/s estimated |
| GeForce RTX 3060, 12GB | $329card only, at launch | 11 GB | 360 GB/s | 11 of 39 | Gemma 4 12B Below every hosted tier | 38 tok/s estimated |
Prices are the card on its own: none of them includes the computer you put it in. Four of the seven are previous-generation cards priced at what they launched at, which is not what you would pay for one today — each card's own page says what to enter instead. Usable memory is what a model gets after the card's own overhead, and speed is on the strongest model in the row, so the column is not a race between equals.
Memory is the gate
A card runs a model or it does not, and nothing about the card changes that except how much memory it has. Qwen3.8 27B is 16 GB of weights before a single token of context, and the cache on top grows with every token you keep. That sum, and not the price, is what puts a model on a card. How the two add up is a page of its own.
It is also where these cards stop. Six of the 34 current models with an index score fit none of the seven cards, and the three strongest open models on the site are among them: GLM-5.3-Flash, Qwen3.8 Flash Next and DeepSeek V4-Flash. The dearest card on the list, 95 GB of memory for $18,000, holds none of them either. For a model that size you are buying unified memory instead: GLM-5.3-Flash runs on the Mac Studio M5 Ultra, 256GB at $10,799, and on nothing cheaper.
Bandwidth is the speed
Once a model fits, the card reads every active weight for every token it writes, so tokens a second tracks memory bandwidth more closely than anything else on the spec sheet. On Qwen3.8 27B the RTX PRO 6000 Blackwell, 96GB does 78 tok/s estimated at 1,792 GB/s, against 30 tok/s estimated for the Radeon AI PRO R9700, 32GB at 640 GB/s. Same model, same context, 2.8× the bandwidth.
Every speed on this page is estimated rather than measured, but not out of thin air: each card's estimate is calibrated against that card's own measured llama.cpp runs, which are listed with their sources on its page. Estimates for small models run high, so read them as an upper bound rather than a promise.
What the same money buys with a computer around it
A card's price here buys the card. You still need the machine it goes in, and this site does not guess at what you would build, so every figure above leaves that cost out. What it can show is what the same money already holds when the computer comes with it. Each card is set against the complete machine nearest it in price, with both prices printed, so a near miss is visible rather than smoothed over.
| Card | Its memory | Models | Nearest complete computer | Its memory | Models |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell, 96GB $18,000, card only | 95 GB | 33 | Mac Studio M5 Ultra, 256GB $10,799 | 192 GB | 38 |
| Radeon AI PRO R9700, 32GB $1,299, card only | 31 GB | 27 | Mac mini M6, 32GB $1,299 | 21 GB | 19 |
| GeForce RTX 5090, 32GB $1,999, card only | 31 GB | 27 | MacBook Pro M5 (14-inch), 16GB $1,999 | 10.5 GB | 10 |
| GeForce RTX 3090, 24GB $1,499, card only | 23 GB | 24 | MacBook Air M5 (15-inch), 16GB $1,499 | 10.5 GB | 10 |
| GeForce RTX 4090, 24GB $1,599, card only | 23 GB | 24 | Mac mini M5 Pro, 24GB $1,699 | 16 GB | 13 |
| GeForce RTX 4080, 16GB $1,199, card only | 15 GB | 13 | Framework Desktop, 32GB $1,269 | 24 GB | 24 |
| GeForce RTX 3060, 12GB $329, card only | 11 GB | 11 | Mac mini M6, 16GB $899 | 10.5 GB | 10 |
On two of the seven the complete computer holds more models than the card does: the Mac Studio M5 Ultra, 256GB holds 38 where the RTX PRO 6000 Blackwell, 96GB holds 33, and the Framework Desktop, 32GB holds 24 where the GeForce RTX 4080, 16GB holds 13. The widest gap is the GeForce RTX 4080, 16GB: 11 models fewer than the Framework Desktop, 32GB, which costs $70 more and is a computer rather than a part for one. What the card has instead is bandwidth: 717 GB/s against 256 GB/s.
Per gigabyte, most of these cards are the expensive way round. The GeForce RTX 3060, 12GB is $30 for each usable gigabyte and the RTX PRO 6000 Blackwell, 96GB is $189. The cheapest gigabyte in a complete computer is the Strix Halo Corsair AI Workstation 300, 64GB at $35, and it arrives with the computer attached. One card beats that — the GeForce RTX 3060, 12GB at $30 — and it is the smallest card here. Divide the price by the usable memory yourself: the two figures are in the tables above. What a card gives back is bandwidth, and only at the top of the range: two of the seven read memory faster than any complete computer here, at 1,792 GB/s against 1,200 GB/s for the Mac Studio M5 Ultra, 96GB. The other five are slower than that machine.
How much use it takes for a card to pay for itself
Everything above is at 500k tokens a day. Pay-back moves with how hard you work a machine, so here is each card at the five levels of use the calculator names, on whichever model it holds that gets there soonest. The levels run from a few chats a day to agents running most of the day.
| Card | Soonest on | 50k a day | 200k a day | 1M a day | 4M a day | 20M a day |
|---|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell, 96GB | Qwen3.8 27B | 2,273 years | 568 years | 114 years | 28 years | 5.7 years |
| Radeon AI PRO R9700, 32GB | Qwen3.8 27B | 167 years | 42 years | 8.3 years | 2.1 years | 5.0 months |
| GeForce RTX 5090, 32GB | Qwen3.8 27B | 254 years | 64 years | 13 years | 3.2 years | 7.6 months |
| GeForce RTX 3090, 24GB | Qwen3.8 27B | 191 years | 48 years | 9.6 years | 2.4 years | 5.7 months |
| GeForce RTX 4090, 24GB | Qwen3.8 27B | 206 years | 51 years | 10 years | 2.6 years | 6.2 months |
| GeForce RTX 4080, 16GB | Ministral 3 14B | 369 years | 92 years | 18 years | 4.6 years | 11 months |
| GeForce RTX 3060, 12GB | Ministral 3 8B | 134 years | 34 years | 6.7 years | 20 months | 4.0 months |
These are the best case for each card and they still run to years at anything short of constant use, because the pay-back is against renting the same model from an API, and the small models a card holds are the cheap ones to rent. The card price is also the whole of the outlay here, so a real build pays back later than the table says. For the same question asked the other way round — which machine pays back soonest at a level of use — see best buys by usage.
Every figure is at 32k context and at the quantisation named on each model's page, with 15:1 input to output, $0.17 per kWh and today's API prices held flat. Models counted are the 39 current ones, the same set every machine page counts. Card prices are list or launch prices for the card alone, read on 2026-09-03, with the source on each card's page. Set two cards against each other in the head-to-heads, rank the models themselves on the leaderboard, or open the calculator and enter what you actually paid.