whichaipc

GPU for local LLM · flagship

NVIDIA GeForce RTX 4090

The fastest 24GB card you'll put in a desktop, if you buy it used and buy it carefully.

VRAM
24 GB
TDP
450 W
Bandwidth
1008 GB/s
£ / GB
£58

What it runs

With 24 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~44B params

Verdict

The same 24GB as a 3090 but much faster and far more efficient, for roughly double the used price.

BEST FOR

  • + fast single-card 24GB inference
  • + dual-card 70B builds
  • + people who feel speed daily

NOT FOR

  • - VRAM-per-pound value hunters
  • - small quiet cases
  • - anyone happy with 3090 speeds

The RTX 4090 is the fastest 24GB card you can drop into a desktop for local LLMs, and I run two of them on my own bench. If you want flagship inference speed without stepping up into workstation-card money, this is the one to hunt for. Used, ideally, now that the 5090 exists and has softened the price.

The take

Buy it for speed layered on top of the same 24GB a 3090 gives you. That’s the whole pitch. You’re paying a hefty premium over a used 3090 for something like 60 to 70 percent more throughput and much better efficiency per token. Worth it if you feel that speed daily; overkill if you don’t. On my bench the pair chews through 70B work a single card can’t touch, and each one stays quicker and cooler per token than the Ampere cards it replaced.

What it’ll actually run

Twenty-four gigabytes runs the same model classes as a 3090: a 32B at Q4 with a generous context window sits comfortably in VRAM, and 8B to 14B models fly. The difference is pace. Expect roughly 50 to 70 tokens/sec on a 13B at Q4 and something like 25 to 40 tokens/sec on a 32B, well ahead of a 3090 and pleasant across long sessions.

Two 4090s give you 48GB, which runs a 70B at Q4 with room for context. That’s my daily setup, and it’s the reason I keep the pair rather than chasing one bigger card. Layers split across the two cards over PCIe with no NVLink needed, and the throughput on a big model still lands somewhere pleasant rather than the crawl you get from CPU offload on a single card. On value, the £/GB-VRAM maths is worse than a 3090. At around £1400 used you’re paying roughly £58 per gigabyte, more than double. You’re buying speed and efficiency here, not cheap VRAM. Software is a non-event: mature Ada drivers, full CUDA, and every stack from Ollama to vLLM targets it first, so setup is boring in the best way.

Who should buy it, and who shouldn’t

Buy it if you want the fastest single-card 24GB inference and you’ll feel the speed every day, or if you’re building a two-card box for 70B work like I have. It’s efficient for its grunt, it ages well, and the driver story is spotless. Tinkerers who’ve outgrown a 3090’s pace are the natural buyers.

Skip it if you mainly care about VRAM per pound; a used 3090 runs the same models for less than half the money, just slower. And if you want new-with-warranty plus 32GB, the 5090 is the current answer at more money again. But for used flagship speed at 24GB, the 4090 is still the card I reach for first.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

Efficient inference

Power limit ~300-350W or undervolt near 950mV (from 450W default)

Owners report around 90 percent of stock throughput at roughly 60 percent power limit; inference is bandwidth bound, so core power can drop with little token/s loss.

Single-card 8B to 13B model

Q4_K_M quant, large context

Runs fully in 24GB at high speed and stays well ahead of a 3090 across long sessions.

Single-card 32B model

Q4_K_M quant, full 24GB

A 32B at Q4_K_M fits with generous context; owners report faster rates than a 3090, roughly 25-40 tokens/sec.

70B model

Two 4090s, 48GB pooled, Q4_K_M, tensor parallel via vLLM

48GB holds a 70B at Q4; there is no NVLink, so tensor parallelism over PCIe is the recommended route for speed.

What owners report

Real first-hand experience gathered from owners and the community.

  • For efficient inference owners report the 4090 is best power limited to roughly 270W to 350W, well under its 450W stock ceiling, as the most efficient point for tokens per watt.

    r/LocalLLaMA

  • Owners running two 4090s for 48GB note the pair runs hot and loud, and recommend vLLM with tensor parallelism over Ollama for markedly faster throughput on big models.

    r/LocalLLaMA

  • A common warning is that 24GB is not enough to hold a 70B model at 4-bit once the context window grows, as attention memory balloons and exhausts VRAM.

    r/LocalLLaMA

  • A large multi-4090 build (eight cards, 192GB) reported around 4,600W real draw under inference load while serving tensor-parallel 70B models, underlining how power-hungry these cards are at scale.

    r/LocalLLM

  • On a single 4090 an owner benchmarked Qwen2.5 32B at about 36 tokens/sec (roughly 86 percent GPU use) and a small Llama 3.2 3B at around 97 tokens/sec using only 8GB of the 24GB, with the card drawing about 277W under load.

    Digital Spaceport

Fact-checked 18 Jul 20269 claims verified against primary sources.
1 claim(s) we couldn't fully verify
  • · typical used price around GBP 1400 / USD 1600 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks used-market pricing; figure is indicative only.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Is a used RTX 4090 worth double a used 3090 for LLMs?+

Only if speed matters to you every day. Both carry 24GB and run the same models, but the 4090 is meaningfully quicker and much more efficient per token. If you sit in front of it for hours, the pace is worth paying for. If you just want the models to run, the 3090 does the same job for less.

Can two RTX 4090s run a 70B model?+

Yes, and it's a lovely setup. Two cards give you 48GB, which holds a 70B at Q4 with room for context. That's the pair I run on my own bench. NVLink isn't available on the 4090, but inference stacks split layers across cards over PCIe perfectly well.

Should I worry about the 12VHPWR power connector?+

Seat it fully until it clicks and route the cable without a tight bend right at the plug, and it's fine. The early melting reports came down to connectors not pushed home properly. Use a good native ATX 3.0 cable rather than a tangle of adapters and you'll have no trouble.