Sunk Cost sunkcost.ai Data checked 2026-09-03

Can a NVIDIA GeForce RTX 4090, 24GB run local LLMs?

Yes — 24 of the 39 open models on this site fit in its 23 GB of usable memory, the strongest being Qwen3.8 27B. Whether that saves you money is a different question, and the answer is usually no.

Price$1,599 at launch — discontinued
Memory24 GB, about 23 GB of it addressable by the GPU at 1008 GB/s
Best model it runsQwen3.8 27B — Sonnet-class, 44 tok/s
Pay-back against the APIPays back in 21 years at 500k tokens a day

Run the numbers on this machine

What it runs

ModelSpeedClassGood atMemory
Qwen3.8 27BQ4_K_M 44 tok/s Sonnet-class 19 GB
Qwen3.6 27BQ4_K_M 43 tok/s Haiku-class 19 GB
Qwen3.6 35B-A3BQ4_K_M 204 tok/s Haiku-class 23 GB
Muse Glimmer 30BQ4_K_M 46 tok/s Haiku-class 18 GB
Gemma 4 26B-A4BQ4_K_M 150 tok/s Haiku-class 18 GB
GLM-4.7-FlashQ4_K_M 145 tok/s Haiku-class 20 GB
Gemma 4 12BQ4_K_M 102 tok/s Below every hosted tier 8.0 GB
Qwen3.5 9BQ4_K_M 121 tok/s Below every hosted tier 6.8 GB
MiniCPM5 2BQ4_K_M 275 tok/s Below every hosted tier 3.0 GB
Qwen3.5 4BQ4_K_M 214 tok/s Below every hosted tier 3.8 GB
Ling 3.0 tinyQ4_K_M 214 tok/s Below every hosted tier 6.4 GB
Granite 4.2 8BQ4_K_M 76 tok/s Below every hosted tier 11 GB

12 more fit; the calculator lists them all.

The specifics

Chip
GeForce RTX 4090 — Ada Lovelace · 24 GB GDDR6X
Memory bandwidth
1008 GB/s
Usable by the GPU
23 GB — Bandwidth: 21 Gbps × 384-bit bus ÷ 8 = 1,008 GB/s (21 Gbps from TechPowerUp, 384-bit from NVIDIA). Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
Power under load
450 W (published) — 450 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. NVIDIA lists 315 W average gaming power. A whole Intel 265K + RTX 4090 PC measured 464-509 W at peak across llama.cpp runs, prompt processing included (Jeff Geerling).
Availability
Launch price. Superseded by the RTX 50 series, so usually bought used for less: enter what you paid.
Sources
source 1, source 2, source 3, source 4, source 5