GPU for local LLM · workstation
NVIDIA RTX PRO 6000 Blackwell
96GB of GDDR7 on one card, the single-GPU ceiling for local AI, priced to match.
- VRAM
- 96 GB
- TDP
- 600 W
- Bandwidth
- 1792 GB/s
- £ / GB
- £78
What it runs
With 96 GB you can comfortably load, at 4-bit quantisation, a dense model up to roughly
~188B params
Verdict
The most VRAM and bandwidth you can put on one card, for people whose work justifies the price.
BEST FOR
- + 70B to 120B on a single card
- + fastest single-GPU inference
- + quiet 96GB workstations
NOT FOR
- - home budgets
- - value-per-pound VRAM
- - anyone a pair of 3090s would satisfy
The RTX PRO 6000 Blackwell is the top of what one card can do for local AI. Ninety-six gigabytes of GDDR7, the same 1792 GB/s bandwidth as a 5090, and a workstation blower to keep it fed, all in two slots. It’s the ceiling. If you’ve ever wanted a model that won’t fit two 3090s to sit on a single GPU, this is the card, and the price says exactly who it’s aimed at.
The take
You’re buying capacity and speed with no compromise, and paying flagship-workstation money for the privilege. Ninety-six gigabytes on one card removes the multi-GPU juggling that big models normally force, and the bandwidth means it’s quick with it. The counterweight is the bill and the 600W draw. For a business where a model running locally earns or saves real money, the maths can work. For a hobby budget, it doesn’t, and that’s fine. This is a tool, priced as one.
What it’ll actually run
Ninety-six gigabytes changes the question from what fits to what you feel like running. A 70B at high quant with a long context, a 120B-class model at Q4, several models loaded at once, or one model with a huge KV cache for long documents. All the things that make a 24GB owner start planning a second card, this does on a single GPU without thinking about it.
Speed keeps pace with the capacity. At 1792 GB/s it sits at the top of single-card token generation, so big models don’t just fit, they run at a usable clip. That combination, large and fast on one card, is the entire reason to choose this over a stack of cheaper GPUs. Two of them pool 192GB for the seriously enormous stuff, though most people who buy one find a single 96GB card already covers the work.
The build is more civilised than the numbers suggest. The blower cooler exhausts out the back, so it stacks and cools better than a triple-slot gaming card, and it takes a single 16-pin. The real demand is the PSU. Six hundred watts of board power wants a large, quality supply with headroom, and a power limit around 450W trims heat and noise at almost no cost to speed.
Who should buy it, and who shouldn’t
Buy it if you need the largest models on one card, want top-tier single-GPU speed, and the work pays the card back. It suits professionals and small teams running big local models daily, where the simplicity of one 96GB card beats managing a multi-GPU rig. If time and tidiness are worth money to you, this is where the money goes.
Don’t buy it on a home budget, because nothing about the price is aimed at hobbyists, and a pair of used 3090s or a single 5090 covers most enthusiast needs for a fraction of the cost. But if 96GB on one fast card is the brief, and the outlay is justified, nothing consumer or prosumer gets you there in a single slot pair like this does.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Big model on one card
A 70B at Q8, or a 120B-class model at Q4, in vLLM or llama.cpp
96GB holds models that otherwise need two or more cards, with room for a long context.
Fast single-GPU inference
Any 24B to 70B fully resident with a large KV cache
At 1792 GB/s token generation is near the top of what a single GPU can do.
Efficient inference (power limit)
nvidia-smi power limit around 450W (from 600W default)
Since generation is bandwidth bound, trimming the cap costs little throughput and cuts heat and noise.
What owners report
Real first-hand experience gathered from owners and the community.
- “
Owners are pairing two of these for 192GB of pooled VRAM to run very large models locally, but several note a single 96GB card already covers most home and prosumer needs.
2 claim(s) we couldn't fully verify
- · indicative price around GBP 7500 / USD 8500 (mid-2026) - No primary source (NVIDIA or TechPowerUp) tracks street pricing; figure is indicative only and varies by reseller.
- · specific 96GB model-fit examples (120B-class at Q4, 70B at Q8) - Illustrative of the capacity; exact fit depends on quant, context and inference stack.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
▶Mukul Tripathi
Inside My 2026 LLM Server: 2x RTX PRO 6000 Blackwell (96GB each)
▶Bijan Bowen
Casual RTX 6000 Pro AI Build - New Local AI Training Setup
- NVIDIA RTX PRO 6000 Blackwell SpecsTechPowerUp · primary source
Common questions
What can 96GB actually run that smaller cards can't?+
A 70B at high quant with a huge context, a 120B-class model at Q4, and long-context work that would spill over on 24GB or 48GB. It also runs multiple models at once, or one model plus a big KV cache, without juggling. This is single-card capacity that otherwise needs two or more GPUs.
Is it worth it over two RTX PRO 6000s or several 3090s?+
Depends what you value. One card at 96GB is simpler, quieter and easier to cool than a stack, and at 1792 GB/s it's fast. Several used 3090s give more total VRAM per pound but need slots, power and cooling, and split models across PCIe. The PRO 6000 is the buy when single-card simplicity and speed are worth the premium.
Can a home PC even power and cool it?+
With planning, yes. It draws 600W over a single 16-pin, so you want a large quality PSU with proper headroom, ideally 1200W or more for the full system. The blower workstation cooler exhausts out the back, which actually makes it easier to cool than a triple-slot gaming card in a stack.