whichaipc

Buyer’s guide

AI servers: what to buy, and when you actually need one

Updated 23 July 2026 · UK prices, checked against live stock

Most people asking what AI server to buy do not need one. That is an awkward thing to open a buyer’s guide with, but I would rather say it than sell you eight grand of hardware you will use at ten per cent. The jump from a desktop with a good graphics card to a proper server is real, and it happens at a specific point. This is about finding that point.

Do you need one at all?

Here is the honest test, and it has nothing to do with how serious you are about this. It is a memory question.

A single 24GB card runs a 32B model at 4-bit comfortably, and on our own bench a 35B mixture-of-experts model returns 133 tokens per second on one card. That is faster than you can read. If that describes what you want to do, a desktop with one good graphics card is the whole answer and a server is a way of spending money.

You cross into server territory when one of three things is true:

  • The model does not fit. A 70B at 4-bit wants around 40GB, which no consumer card has. This is the common one.
  • Several people use it. Concurrency is where the big cards pull away - four parallel requests on our 80B run aggregate to 315 tokens per second.
  • It runs unattended, all the time. Different problem: thermals, noise and power over months rather than peak speed.

Wanting to run the biggest model you have heard of is not on that list. The most useful finding from our own benchmarking is that architecture beats size - an 8B model tops our table at 176 tokens per second while a dense 27B manages 61. You may need far less machine than you think.

What actually makes it a server

The word covers three different things people are shopping for, and they have almost nothing in common:

  • A workstation - a tower under a desk with one or two big professional cards. Quiet enough to sit next to. This is what most people asking the question want.
  • A rack unit - 2U or 4U, passive cards, screaming fans, lives in a cupboard or a colo. Denser and cheaper per GB of memory, unpleasant anywhere near people.
  • A repurposed desktop - an old tower with a couple of used cards in it. Perfectly good, and where a lot of good home setups actually start.

The distinction that matters when you are buying is cooling, not form factor. Professional cards come in active and passive versions and the passive ones have no fan at all, because they assume a chassis blowing air through them at a rate no tower case manages. Put a passive card in a normal case and it will throttle, then it will fail. The Server Edition in the table below is the one to watch: it is passive. Max-Q is a different thing again, a lower-power variant that still has its own cooler.

The tiers, and what they cost in the UK

Prices below are scraped from Overclockers UK rather than typed in, so they move with stock. Everything here is the card, not a complete machine - add roughly £1,200 to £2,000 for a chassis, CPU, memory and a power supply that can take it.

Card Memory UK price Stock
H200 NVL Passive 141GB GDDR6 Server Graphics Card 141GB £30,000 Pre order Check price →
RTX PRO 6000 Blackwell Server Edition 96GB GDDR7 96GB £13,500 In stock Check price →
RTX PRO 6000 Blackwell Workstation Edition 96GB GDDR7 96GB £12,000 In stock Check price →
RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB GDDR7 96GB £12,000 Pre order Check price →
RTX PRO 5000 Blackwell Workstation Edition 72GB GDDR7 (OEM) 72GB £8,000 In stock Check price →
RTX 6000 Ada 48GB GDDR6 PCI-Express 48GB £7,500 Pre order Check price →

Source: overclockers.co.uk, captured 2026-07-22. Affiliate links - it costs you nothing and we get a small cut.

Reading that table as a ladder:

The 48GB tier is the first real step past consumer hardware, and it is the one that gets 70B models running. The RTX 6000 Ada at around £7,500 is the new option. It is also where a used card makes the most sense, which I will come back to.

The 96GB tier is the RTX PRO 6000 Blackwell, and at roughly £12,000 for the workstation version it is the sweet spot for a serious single-card build. One card, 96GB, no tensor-parallel complexity, no peer-to-peer problems, active cooling. If somebody asked me to specify a machine to run large models properly and quietly, it would be built around this.

The 141GB tier is the H200 NVL at around £30,000, and it is not really a home purchase. If you are looking at it you are buying for a business, and the calculation is against renting rather than against a cheaper card.

Worth saying plainly: two 48GB cards give you 96GB for less than one 96GB card, and our own 80B run does exactly that across two modded 4090s at 118 tokens per second. It works. It also costs you a fussier setup - peer-to-peer disabled, PCIe lanes to think about, more heat in the same box. Sometimes the single card is worth the premium for the evening you do not spend debugging it.

Why VRAM decides almost everything

Generating a token means reading the model out of memory. All of it, every time. So two things follow, and they are the whole of hardware selection for inference.

First, the model has to fit, or it spills into system RAM and the speed collapses by an order of magnitude. Not degrades - collapses. Dual-channel DDR5 gives you roughly 96GB/s against a 4090’s 1,000GB/s.

Second, once it fits, memory bandwidth sets the ceiling and nothing beats it. That is why a card’s bandwidth figure predicts its speed better than any compute number on the box.

The rough sum for what fits: half the parameter count in GB at 4-bit, plus 2-4GB of headroom. A 70B needs about 40GB. A 32B needs about 18GB. Our VRAM calculator does it properly, including context length, which grows the requirement more than people expect on long prompts.

The used enterprise route

The tier above is what things cost new. There is a much cheaper path, and for a home machine it is often the right one.

Used RTX A6000s (48GB) and the previous generation of data-centre cards turn up regularly at a fraction of new prices, because businesses cycle hardware on a schedule rather than when it stops being useful. A 48GB card that cost £5,000 new and comes off a three-year lease is a different proposition entirely.

What to check before you buy: that it is the active-cooled variant unless you have a chassis for it, that the seller will let you run it for an hour before the return window closes, and what the card was doing. Cards from render farms and inference clusters have had an easy life at steady load. The ones to be wary of are the ones with no history at all.

Our GPU comparison covers the consumer end of this, including the used 3090 route that remains the best value in the whole market if 24GB is enough.

What catches people out

In rough order of how often it ruins somebody’s week:

  • Passive cards in a normal case. Covered above and worth repeating, because the Server Edition often looks like the bargain in a price list. It is not cheaper, it is a different product.
  • Noise. A 2U chassis at load is hard to sit in a room with. People buy rack gear for the price per GB and then discover they cannot work next to it.
  • Power and the wall. A single big card is fine. Two 350W cards plus a CPU under sustained load is a real draw on a UK ring main, and the sustained part matters - inference is not a gaming burst, it is hours at full load. Our own testing found 330W performed identically to 370W, so power limiting costs you nothing and is worth doing.
  • PCIe lanes. Consumer boards split to x8/x8 for two cards, which is fine for inference, but check the second slot is not x4 and that an M.2 drive is not stealing the lanes.
  • Buying for training when you mean inference. They have different requirements. Almost nobody running a home setup is training anything, and specifying for it is expensive.

If you are still deciding between one big card and two smaller ones, the measured benchmarks are the most useful thing on this site - real numbers from our own rig, with the configuration that produced each one.

Share