Can a NVIDIA GeForce RTX 4090, 24GB run local LLMs?
Yes — 24 of the 39 open models on this site fit in its 23 GB of usable memory, the strongest being Qwen3.8 27B. Whether that saves you money is a different question, and the answer is usually no.
Price$1,599 at launch — discontinued
Memory24 GB, about 23 GB of it addressable by the GPU at 1008 GB/s
Best model it runsQwen3.8 27B — Sonnet-class, 44 tok/s
Pay-back against the APIPays back in 21 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Qwen3.8 27BQ4_K_M | 44 tok/s | Sonnet-class | 19 GB | |
| Qwen3.6 27BQ4_K_M | 43 tok/s | Haiku-class | 19 GB | |
| Qwen3.6 35B-A3BQ4_K_M | 204 tok/s | Haiku-class | 23 GB | |
| Muse Glimmer 30BQ4_K_M | 46 tok/s | Haiku-class | 18 GB | |
| Gemma 4 26B-A4BQ4_K_M | 150 tok/s | Haiku-class | 18 GB | |
| GLM-4.7-FlashQ4_K_M | 145 tok/s | Haiku-class | 20 GB | |
| Gemma 4 12BQ4_K_M | 102 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 121 tok/s | Below every hosted tier | 6.8 GB | |
| MiniCPM5 2BQ4_K_M | 275 tok/s | Below every hosted tier | 3.0 GB | |
| Qwen3.5 4BQ4_K_M | 214 tok/s | Below every hosted tier | 3.8 GB | |
| Ling 3.0 tinyQ4_K_M | 214 tok/s | Below every hosted tier | 6.4 GB | |
| Granite 4.2 8BQ4_K_M | 76 tok/s | Below every hosted tier | 11 GB |
12 more fit; the calculator lists them all.
The specifics
- Chip
- GeForce RTX 4090 — Ada Lovelace · 24 GB GDDR6X
- Memory bandwidth
- 1008 GB/s
- Usable by the GPU
- 23 GB — Bandwidth: 21 Gbps × 384-bit bus ÷ 8 = 1,008 GB/s (21 Gbps from TechPowerUp, 384-bit from NVIDIA). Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
- Power under load
- 450 W (published) — 450 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. NVIDIA lists 315 W average gaming power. A whole Intel 265K + RTX 4090 PC measured 464-509 W at peak across llama.cpp runs, prompt processing included (Jeff Geerling).
- Availability
- Launch price. Superseded by the RTX 50 series, so usually bought used for less: enter what you paid.
- Sources
- source 1, source 2, source 3, source 4, source 5