Can a NVIDIA GeForce RTX 4080, 16GB run local LLMs?
Yes — 13 of the 39 open models on this site fit in its 15 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.
Price$1,199 at launch — discontinued
Memory16 GB, about 15 GB of it addressable by the GPU at 716.8 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 74 tok/s
Pay-back against the APIPays back in 93 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Gemma 4 12BQ4_K_M | 74 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 87 tok/s | Below every hosted tier | 6.8 GB | |
| MiniCPM5 2BQ4_K_M | 198 tok/s | Below every hosted tier | 3.0 GB | |
| Qwen3.5 4BQ4_K_M | 154 tok/s | Below every hosted tier | 3.8 GB | |
| Ling 3.0 tinyQ4_K_M | 125 tok/s | Below every hosted tier | 6.4 GB | |
| Granite 4.2 8BQ4_K_M | 55 tok/s | Below every hosted tier | 11 GB | |
| gpt-oss-20bMXFP4 | 103 tok/s | Below every hosted tier | 13 GB | |
| Gemma 4 E4BQAT Q4_0 | 87 tok/s | Below every hosted tier | 5.7 GB | |
| LFM2.5 2.6BQ4_K_M | 266 tok/s | Below every hosted tier | 2.2 GB | |
| Ministral 3 14BQ4_K_M | 43 tok/s | Below every hosted tier | 14 GB | |
| Ministral 3 8BQ4_K_M | 60 tok/s | Below every hosted tier | 9.8 GB | |
| Ornith 1.5 9BQ4_K_M | 86 tok/s | not yet placed | 6.9 GB |
1 more fit; the calculator lists them all.
The specifics
- Chip
- GeForce RTX 4080 — Ada Lovelace · 16 GB GDDR6X
- Memory bandwidth
- 716.8 GB/s
- Usable by the GPU
- 15 GB — Bandwidth: 22.4 Gbps × 256-bit bus ÷ 8 = 716.8 GB/s (22.4 Gbps from TechPowerUp, 256-bit from NVIDIA). The 4080 SUPER has the same 16 GB and bus but 23 Gbps memory (736 GB/s) and a different price, so it is not this entry. Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
- Power under load
- 320 W (published) — 320 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. NVIDIA lists 251 W average gaming power.
- Availability
- Launch price. Replaced by the $999 RTX 4080 SUPER and then the RTX 50 series.
- Sources
- source 1, source 2, source 3, source 4