whichaipc

GPU for local LLM · budget

NVIDIA GeForce RTX 3060 12GB

The cheapest sensible way into local LLMs, thanks to an oddly generous 12GB.

VRAM
12 GB
TDP
170 W
Bandwidth
360 GB/s
£ / GB
£19

What it runs

With 12 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~20B params

Verdict

12GB for around £230 makes this the floor for running local models properly.

BEST FOR

  • + first local-AI build
  • + 7B to 14B models
  • + low-power always-on setups

NOT FOR

  • - 30B-plus models
  • - fast long-context work
  • - anyone who'll outgrow 12GB in months

If you’re curious about running LLMs at home but you’re not ready to spend real money, the RTX 3060 12GB is the floor. It’s the cheapest card that still does the job properly, and the one to point a first-timer at. Not fast, but properly capable, and cheap enough that a mistake doesn’t sting.

The take

This is an entry ticket, not a performance card. What makes it work is that oddly generous 12GB of VRAM; NVIDIA fitted more memory here than on some pricier Ampere cards, and for local AI that’s exactly the spec that counts. At roughly £230 used, it’s the least you can spend and still run models worth running. Slow-ish, cool, and cheap. That’s the deal.

What it’ll actually run

Twelve gigabytes lands you squarely in 7B to 14B territory, which covers a lot of real work. A 7B or 8B model at Q4 sits fully in VRAM with headroom for context, and a 14B at Q4 fits if you keep the context modest. That’s enough for chat, coding help, summarising, and most assistant tasks.

Speeds are where the budget shows. Expect something like 20 to 35 tokens/sec on a 7B at Q4, dropping toward 10 to 18 tokens/sec on a 14B. Fine for interactive use, less fun for long documents or heavy batch work. Quick enough to read along with, slow enough that you’ll notice on anything meaty. The 360 GB/s bandwidth is the ceiling, and it’s why bigger models with offload crawl; keep everything in VRAM and it stays pleasant.

On value, £230 for 12GB is around £19 per gigabyte, the best £/GB-VRAM figure of the three cards here. You’re paying less per gigabyte than a 3090, you’re just getting fewer of them. For a starter box, that’s the right trade. And because it’s plain Ampere on mature CUDA drivers, everything from Ollama to llama.cpp just works with no fuss.

Who should buy it, and who shouldn’t

Buy it if this is your first local-AI machine, or if you want a low-power card that can sit running an assistant all day without heating the room. It’s forgiving to build with: two slots, one 8-pin, 170W, and it drops into almost any case. Perfect for learning what local models can do before committing more.

Don’t buy it if you already know you want 30B-plus models or fast long-context work; you’ll bump into the 12GB wall quickly and wish you’d bought a 3090 for the extra headroom. The step up is real money, so think about where you’ll be in six months. But if the question is simply “what’s the cheapest card that runs local LLMs properly”, the 12GB 3060 is the sensible answer, and the 8GB version is not.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

First local-AI box

Q4_K_M 7B to 8B, full GPU offload

Sits fully in 12GB with headroom for context; one 8-pin, 170W, drops into almost any case.

14B ceiling

Q4_K_M, modest context

A 14B at Q4 fits if you keep the context small; the 360 GB/s bus is the limiter.

Always-on assistant

leave running at 170W board power

Low idle draw makes it the easiest card here to run an assistant all day.

What owners report

Real first-hand experience gathered from owners and the community.

  • Testing a dual 3060 12GB rig (24GB pooled) on Ollama, Digital Spaceport measured Gemma 3 12B Q8 at about 19.7 tokens/sec and Deepcoder 14B Q8 at around 17, but big models starved on prompt processing: Gemma 3 27B Q4 fell to roughly 6 tokens/sec and QwQ 32B Q4 to about 12.

    Digital Spaceport

  • On a single card, owners report a 3060 12GB running 8B models at Q4 around 42 tokens/sec, quick enough for interactive chat, while a 32B with offload crawls at roughly 3 to 5 tokens/sec.

    modelfit.io / r/LocalLLaMA

Fact-checked 18 Jul 20266 claims verified against primary sources.
1 claim(s) we couldn't fully verify
  • · typical used price around GBP 230 / USD 280 (mid-2026) - No primary source tracks used-market pricing; community listings cluster near 200 to 250 dollars, figure is indicative only.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Why the 12GB 3060 and not the 8GB version?+

The 12GB card is the whole point. NVIDIA gave the base 3060 more VRAM than the tier above it, which is odd for gaming but brilliant for LLMs. The 8GB 3060 exists and should be avoided for AI; those extra 4GB are the difference between running a useful model and constantly hitting limits.

What's the biggest model an RTX 3060 12GB can run?+

A 14B model at Q4 fits with a modest context window, and 7B to 8B models run comfortably with room to spare. Anything from 30B up needs offload to system RAM, which the 3060's 360 GB/s bandwidth makes slow enough that you won't enjoy it.

Is 170W low enough to leave running all day?+

Yes. At 170W board power and far less at idle, the 3060 is the easiest card here to run as an always-on local assistant. It fits two slots, needs one 8-pin, and won't cook a small case.