whichaipc

GPU for local LLM · flagship

NVIDIA GeForce RTX 5090

Fastest consumer card for local AI, with 32GB, if the price and 575W don't put you off.

VRAM
32 GB
TDP
575 W
Bandwidth
1792 GB/s
£ / GB
£59

What it runs

With 32 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~60B params

Verdict

The quickest way to run bigger models on one card, priced so it only makes sense if speed pays you back.

BEST FOR

  • + fast single-card inference
  • + 32B models with big context
  • + new-with-warranty buyers

NOT FOR

  • - value-per-pound builds
  • - low-power always-on rigs
  • - anyone happy at 30 tokens/sec

If you want the quickest single card for running LLMs at home and budget isn’t the deciding factor, the RTX 5090 is the one. Thirty-two gigabytes of GDDR7 and the highest memory bandwidth on any consumer card make it the speed pick. It’s for people whose time on tokens is worth real money.

The take

You’re paying a flagship premium for two things: speed and 32GB. Both are real. At 1792 GB/s the 5090 roughly doubles a 3090’s memory bandwidth, and bandwidth is what sets token generation speed, so it feels quicker on the same model. The catch is the bill, around £1,900, and the 575W it pulls to get there. Fast and future-proof, but neither cheap nor gentle on your power draw.

What it’ll actually run

Thirty-two gigabytes is a meaningful step past 24GB. You’ll run a 32B model at Q4 with a generous context window fully in VRAM, or push a 32B to higher quant for better quality, none of which a 24GB card manages without offload. Smaller 7B to 14B models run so far inside the card’s limits they respond near-instantly.

Speed is the headline. On a 32B at Q4 you’re looking at token rates comfortably ahead of anything Ampere can do, and on smaller models it’s well into the hundreds of tokens/sec range for short generations. If you batch work, feed it long documents, or just hate waiting, that throughput is the reason to buy.

The value lens is less kind. At around £1,900 for 32GB that’s roughly £59 per gigabyte, more than double a 3090’s £/GB-VRAM. You’re not buying cheap capacity here; you’re buying speed and the newest architecture. Judge it on whether faster tokens actually pay you back. One caveat worth knowing: Blackwell is newer, so the odd inference build still needs an up-to-date CUDA toolkit to behave, where Ampere just runs.

Who should buy it, and who shouldn’t

Buy it if you want the fastest single-card local AI going, need 32GB in one slot’s worth of VRAM budget, and value a new card with a warranty over used-market gambling. It suits people running models as part of paid work, where waiting less is worth the outlay. Two slots and a stock 12VHPWR keep the build tidy, provided your PSU is up to 575W.

Don’t buy it for value; the maths favours a 3090, or two, every time on pure VRAM per pound. Skip it too if you want a low-power always-on box, because 575W is the opposite of that. If you’d be just as happy at 30 tokens/sec on a 32B, save the difference. But if speed and 32GB on one current card are the brief, nothing consumer beats it.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

Fast single-card inference (32B)

Qwen2.5-Coder 32B Q4_K_M (~20GB) in LM Studio

One owner measured around 62 tokens/sec with roughly 22GB of the 32GB VRAM in use; the 8-bit quant at about 34GB will not fit.

Small model, high throughput

DeepSeek-R1-Distill-Llama 8B in LM Studio

Same tester measured around 202 tokens/sec, averaging under 300W of power.

Efficient inference (power limit)

nvidia-smi power limit 400W (from 575W default)

Dropping the cap from 575W to 400W cost only a few tokens/sec on a 32B (62 down to 58), since generation is bandwidth bound, not core-power bound.

70B model (single card)

Llama 3.1 70B Q4 with partial offload

Community benchmarks put a 70B Q4 near 14 tokens/sec on one 5090; 32GB cannot hold it fully, so speed depends on offload.

What owners report

Real first-hand experience gathered from owners and the community.

  • Running Qwen2.5-Coder 32B at Q4 in LM Studio, an owner measured about 62 tokens/sec on the desktop 5090 with roughly 22GB of the 32GB used; the 8-bit version at about 34GB would not fit.

    Alex Ziskind

  • The same tester saw the card pull up to around 570W of its 600W cap on a long prompt, but capping power at 400W only dropped a 32B from 62 to 58 tokens/sec, because token generation is memory-bandwidth bound rather than core-power bound.

    Alex Ziskind

  • Early adopters comparing the 5090 to a 4090 on the same models reported roughly 60 to 80 percent more tokens per second, attributing it to about 30 percent more and 30 percent faster memory.

    r/LocalLLaMA

Fact-checked 18 Jul 20269 claims verified against primary sources.
2 claim(s) we couldn't fully verify
  • · indicative price around GBP 1900 / USD 2000 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks street pricing; figure is indicative only.
  • · single-card tokens/sec figures (62 t/s on 32B Q4, 202 t/s on 8B, roughly 14 t/s on 70B) - Community and owner benchmarks; vary with quant, context, driver and inference stack.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Is the RTX 5090's 32GB worth it over a 24GB 3090?+

The extra 8GB lets you run a 32B model with a much larger context, or a 32B at higher quant, without offload. Whether that's worth roughly three times the price depends on whether you need the speed and the headroom daily. For most home users, two 3090s give more total VRAM for less.

How much faster is the 5090 for LLMs?+

A lot, at the memory-bandwidth level that governs token generation. At 1792 GB/s it roughly doubles the 3090's bandwidth, so expect token rates well ahead of any Ampere card on the same model. If throughput is the bottleneck in your workflow, this is where the money goes.

Do I need a new power supply for a 575W card?+

Probably. At 575W board power over a 12VHPWR connector, you want a quality modern PSU with the native cable and plenty of headroom, ideally 1000W or more for a full system. Seat the connector fully; the earlier 12VHPWR reputation came largely from partial insertion.