GPU for local LLM · value
NVIDIA GeForce RTX 3080
Fast, cheap used, and boxed in by 10GB of VRAM for local models.
- VRAM
- 10 GB
- TDP
- 320 W
- Bandwidth
- 760 GB/s
- £ / GB
- £35
What it runs
With 10 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly
~16B params
Verdict
Properly fast small-model performance on a budget, undone by a 10GB ceiling you'll hit quickly.
BEST FOR
- + quick 7B to 8B models
- + repurposed gaming cards
- + budget speed
NOT FOR
- - 30B-plus models
- - long-context work
- - anyone who'll want more headroom soon
The RTX 3080 is fast, cheap on the used market, and hobbled for local LLMs by one number: 10GB. It’ll run smaller models quickly and happily, but the memory ceiling arrives sooner than you’d like. A good card fighting a VRAM problem, and worth understanding before you buy.
The take
Buy it if you’re mostly running 7B to 8B models and you want them quick for not much money. The GA102 silicon and 760 GB/s of GDDR6X are properly fast, well ahead of a 3060 by a wide margin. The trouble is that 10GB. It’s enough for the small end and not much more, and for local AI, VRAM is usually the wall you hit first. If your models fit, it’s brilliant value. If they don’t, no amount of speed rescues you.
What it’ll actually run
Ten gigabytes is a real constraint. A 7B or 8B model at Q4 fits with a sensible context and runs fast. A 13B at Q4 is a squeeze; you can manage it with a trimmed context, but you’re juggling memory rather than relaxing. Anything from 30B up is off the table without heavy offload, and offload on a card this quick feels like a waste of good silicon.
Where models do fit, it flies. Expect something like 45 to 65 tokens/sec on a 7B at Q4, comfortably ahead of budget Ampere and quick enough that you’ll never wait on it. On value, around £350 for 10GB works out near £35 per gigabyte, which isn’t the bargain a 3090 is; you’re paying for the speed and the fast memory, not for capacity. It’s plain Ampere on mature CUDA drivers, so Ollama, llama.cpp and the rest set up without a fight.
Who should buy it, and who shouldn’t
Buy it if you know you’ll live in the 7B to 8B range and you want those models snappy on a budget, or if you already have one from a gaming build and fancy seeing what it does. It’s fast, forgiving, and a fine way to learn what local models can do. For the money, few cards feel this quick on small models.
Don’t buy it if you want room to grow; the 10GB ceiling is close, and you’ll feel it the moment you reach for a 32B. For a bit more, a used 3090 more than doubles your VRAM and opens up the models this card simply can’t hold. Where will you be in six months? The gap between a card that fits your model and one that doesn’t is night and day. As a fast small-model card the 3080 delivers well; as a do-it-all local box, it runs out of memory long before it runs out of speed.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Fast small model
Q4_K_M 7B to 8B, full GPU offload
Fits inside 10GB with room for a sensible context and runs well above reading speed.
Squeeze a 13B
Q4_K_M, trimmed context
A 13B at Q4 is tight in 10GB; you manage the KV cache rather than relax.
Modern small quants
aggressive 4-bit quant with flash attention
Newer small models (Gemma-class) at 4-bit stretch what 10GB can do; flash attention works on Ampere.
What owners report
Real first-hand experience gathered from owners and the community.
- “
A used 3080 sits around 400 dollars and its 10GB comfortably holds aggressive 4-bit quants of modern small models; VRAM is the ceiling here, not compute. The die is the same as the A5000, and a later 12GB variant eases the limit slightly.
- “
On a 7B at Q4 the GA102 silicon and 760 GB/s of GDDR6X are quick; owners report small models streaming faster than you can read, with the 10GB wall arriving before the speed ever does.
1 claim(s) we couldn't fully verify
- · typical used price around GBP 350 / USD 400 (mid-2026) - No primary source tracks used-market pricing; Ai Flux cites roughly 400 dollars (April 2026), figure is indicative only.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
- Structured comparison of hardware tokens per second (incl. RTX 3080 10GB)r/LocalLLaMA · thread
- GeForce RTX 3080 specificationsTechPowerUp · primary source
Common questions
Is 10GB enough for running local LLMs?+
Just about, for the small end. A 7B or 8B model at Q4 fits with a sensible context and runs fast. A 13B is a tight squeeze with a trimmed context, and anything bigger needs offload that spoils the card's speed. If you know you'll stay small, 10GB works; if not, it's the wall you'll hit first.
How does the RTX 3080 compare to the 3060 12GB for AI?+
The 3080 is much faster and has more bandwidth, but the 3060 has more VRAM. For small models that fit in 10GB the 3080 wins on speed easily. For anything that needs 12GB, the humble 3060 quietly does what the 3080 can't.
Should I buy a 3080 or save for a 3090?+
If you can stretch, the 3090. Its 24GB more than doubles what you can run and opens up 32B-class models the 3080 can't hold. The 3080 only makes sense if you're certain you'll live in the small-model range and want the speed cheaply.
