Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a NVIDIA GeForce RTX 4080, 16GB run local LLMs?

Yes — 13 of the 39 open models on this site fit in its 15 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.

Price$1,199 at launch — discontinued
Memory16 GB, about 15 GB of it addressable by the GPU at 716.8 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 74 tok/s
Pay-back against the APIPays back in 93 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
Gemma 4 12BQ4_K_M 74 tok/s Below every hosted tier 8.0 GB
Qwen3.5 9BQ4_K_M 87 tok/s Below every hosted tier 6.8 GB
MiniCPM5 2BQ4_K_M 198 tok/s Below every hosted tier 3.0 GB
Qwen3.5 4BQ4_K_M 154 tok/s Below every hosted tier 3.8 GB
Ling 3.0 tinyQ4_K_M 125 tok/s Below every hosted tier 6.4 GB
Granite 4.2 8BQ4_K_M 55 tok/s Below every hosted tier 11 GB
gpt-oss-20bMXFP4 103 tok/s Below every hosted tier 13 GB
Gemma 4 E4BQAT Q4_0 87 tok/s Below every hosted tier 5.7 GB
LFM2.5 2.6BQ4_K_M 266 tok/s Below every hosted tier 2.2 GB
Ministral 3 14BQ4_K_M 43 tok/s Below every hosted tier 14 GB
Ministral 3 8BQ4_K_M 60 tok/s Below every hosted tier 9.8 GB
Ornith 1.5 9BQ4_K_M 86 tok/s not yet placed 6.9 GB

1 more fit; the calculator lists them all.

The specifics

Chip
GeForce RTX 4080 — Ada Lovelace · 16 GB GDDR6X
Memory bandwidth
716.8 GB/s
Usable by the GPU
15 GB — Bandwidth: 22.4 Gbps × 256-bit bus ÷ 8 = 716.8 GB/s (22.4 Gbps from TechPowerUp, 256-bit from NVIDIA). The 4080 SUPER has the same 16 GB and bus but 23 Gbps memory (736 GB/s) and a different price, so it is not this entry. Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
Power under load
320 W (published) — 320 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. NVIDIA lists 251 W average gaming power.
Availability
Launch price. Replaced by the $999 RTX 4080 SUPER and then the RTX 50 series.
Sources
source 1, source 2, source 3, source 4