whichaipc

GPU for local LLM · workstation

NVIDIA RTX A6000

48GB of blower-cooled Ampere in two slots, the used workstation card that runs big models on a single GPU.

VRAM
48 GB
TDP
300 W
Bandwidth
768 GB/s
£ / GB
£67

What it runs

With 48 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly

~92B params

Verdict

48GB on one card, in two slots, at 300W, if you can stomach used workstation pricing.

BEST FOR

  • + 70B on a single card
  • + quiet two-slot blower builds
  • + ECC and pro driver stability

NOT FOR

  • - value-per-pound VRAM
  • - fastest token speed
  • - tight budgets

The RTX A6000 is the workstation card that keeps turning up on home AI benches, and for one reason. Forty-eight gigabytes of memory, on a single dual-slot card, at 300W, cooled by a blower that dumps its heat out the back. It’s not fast and it isn’t cheap, but it runs a 70B on one GPU in a build you can actually keep quiet. That’s a rare combination.

The take

You’re paying for packaging, not speed. Two 3090s pool the same 48GB for a good deal less money, so on pure value the A6000 loses. What it gives you instead is 48GB in two slots at 300W with ECC memory and pro drivers, which makes for a clean, stable, single-card machine. If a tidy build and single-GPU simplicity are worth real money to you, that’s the case. If they aren’t, dual 3090s win.

What it’ll actually run

Forty-eight gigabytes changes what one card can hold. A 70B at Q4 fits with room for context, and you don’t need a second GPU or CPU offload to get there. That alone is the reason people buy it. Below that, a 32B runs at higher quant and much longer context than any 24GB card manages, which helps with document work and long chats.

Speed is the compromise. At 768 GB/s the bandwidth trails a 3090, so tokens come out at a steady rather than snappy pace, and for the price that stings if you were expecting quick. Judge this card on capacity per slot and per watt, not on throughput. It’s a fit-it-in-one-card tool.

The build side is where it shines. The blower cooler self-exhausts, so two of them stack in adjacent slots and stay sane on air, where two triple-slot gaming cards would choke. Three hundred watts and a single power input keep the wiring simple. For a quiet, always-on box with a lot of VRAM, little else is this civilised.

Who should buy it, and who shouldn’t

Buy it if you want a large model on one card in a clean, quiet, workstation-style build, and 48GB of ECC in two slots is worth paying over. It suits people who value stability and packaging, or who plan to run two cards for 96GB on air. Used pricing is the way in, so buy from a seller who will stand behind it.

Skip it if value is the priority, because dual 3090s give the same 48GB for less. Skip it too if you want the fastest tokens per pound, where newer cards do better. But for 48GB in a two-slot, 300W package that stays quiet, the A6000 still earns its keep.

Settings people actually run

The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.

70B on one card

Llama 3.1 70B at Q4_K_M in llama.cpp or vLLM

A 70B at Q4 fits inside 48GB with usable context, no second card or offload needed.

Large context on a 32B

A 32B at Q5 to Q8 with a long context window

The 48GB headroom lets you run higher quant and much longer context than a 24GB card allows.

Two-card scaling

Two A6000s, 96GB pooled

Two blowers stack in adjacent slots and stay manageable on air, which dual triple-slot gaming cards rarely do.

Fact-checked 19 Jul 20269 claims verified against primary sources.
1 claim(s) we couldn't fully verify
  • · typical used price around GBP 3200 / USD 4000 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks used-market pricing; figure is indicative only.

Hands-on reviews we drew on

We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.

Common questions

Why buy an A6000 over two 3090s for the same 48GB?+

Packaging and power. The A6000 gives you 48GB on one card, in two slots, at 300W, with a blower that exhausts out the back and ECC memory. Two 3090s cost less and pool the same 48GB, but they need more slots, more power, and more cooling. You pay the A6000 premium for a clean, single-card, workstation-grade build.

Is it fast for LLMs?+

Adequate, not quick. At 768 GB/s its bandwidth sits below a 3090, so token speed is modest for the money. The point of this card isn't speed, it's fitting a large model in one card's VRAM. If throughput matters more than capacity, newer cards do better per pound.

Is the A6000 the same as the RTX 6000 Ada?+

No. The A6000 is Ampere with 48GB of GDDR6. The RTX 6000 Ada is the newer Ada-generation card, also 48GB but faster and pricier. This page is the Ampere A6000, which is the one that shows up used at a more approachable price.