Can a NVIDIA GeForce RTX 3060 12GB, 12GB run local LLMs?
Yes — 11 of the 39 open models on this site fit in its 11 GB of usable memory, the strongest being Gemma 4 12B. Whether that saves you money is a different question, and the answer is usually no.
Price$329 at launch — discontinued
Memory12 GB, about 11 GB of it addressable by the GPU at 360 GB/s
Best model it runsGemma 4 12B — Below every hosted tier, 38 tok/s
Pay-back against the APIPays back in 26 years at 500k tokens a day
Run the numbers on this machine
What it runs
| Model | Speed | Class | Good at | Memory |
|---|---|---|---|---|
| Gemma 4 12BQ4_K_M | 38 tok/s | Below every hosted tier | 8.0 GB | |
| Qwen3.5 9BQ4_K_M | 45 tok/s | Below every hosted tier | 6.8 GB | |
| MiniCPM5 2BQ4_K_M | 103 tok/s | Below every hosted tier | 3.0 GB | |
| Qwen3.5 4BQ4_K_M | 80 tok/s | Below every hosted tier | 3.8 GB | |
| Ling 3.0 tinyQ4_K_M | 45 tok/s | Below every hosted tier | 6.4 GB | |
| Granite 4.2 8BQ4_K_M | 29 tok/s | Below every hosted tier | 11 GB | |
| Gemma 4 E4BQAT Q4_0 | 31 tok/s | Below every hosted tier | 5.7 GB | |
| LFM2.5 2.6BQ4_K_M | 139 tok/s | Below every hosted tier | 2.2 GB | |
| Ministral 3 8BQ4_K_M | 31 tok/s | Below every hosted tier | 9.8 GB | |
| Ornith 1.5 9BQ4_K_M | 45 tok/s | not yet placed | 6.9 GB | |
| Spark-X2.5 4BQ4_K_M | 79 tok/s | not yet placed | 3.9 GB |
The specifics
- Chip
- GeForce RTX 3060 12GB — Ampere · 12 GB GDDR6
- Memory bandwidth
- 360 GB/s
- Usable by the GPU
- 11 GB — Bandwidth: 15 Gbps × 192-bit bus ÷ 8 = 360 GB/s (15 Gbps from TechPowerUp, 192-bit from NVIDIA). Usable memory is the VRAM less 1 GB: the margin llama.cpp's automatic fitting (--fit) leaves free on each device by default (common/common.h, fit_params_target = 1 GiB). Treated as nominal GB, like the other machines here.
- Power under load
- 170 W (published) — 170 W is the maker's rated board power for the card alone. The rest of the PC adds more, and LLM decoding usually draws less than the full rating. A whole Intel 265K + RTX 3060 PC measured 214-229 W at peak during llama.cpp runs, prompt processing included (Jeff Geerling).
- Availability
- Launch price. Two generations old; usually a used or clearance purchase.
- Sources
- source 1, source 2, source 3, source 4, source 5