AI PC · desktop
Framework Desktop (Ryzen AI Max+ 395)
128GB of unified memory on plain x86, for roughly half what the rivals cost.
- Memory
- 128 GB
- Bandwidth
- 256 GB/s
- AI compute
- 50 TOPS
- £ / GB
- £13
What it runs
With 128 GB of memory you can load, at 4-bit, a model up to roughly
~252B params
Verdict
The value pick of the unified-memory machines: same 128GB, plain x86, half the price.
BEST FOR
- + most model per pound
- + people who want Windows or Linux
NOT FOR
- - raw speed
- - anyone who needs CUDA
The Framework Desktop takes the interesting bit of the Strix Halo chip - up to 128GB of memory the graphics side can treat as its own - and drops it into a proper little x86 desktop you can actually tinker with. For anyone who wants big-model headroom without leaving Windows or Linux behind, this is the friendly option.
The take
This is the value pick of the unified-memory crowd, and it’s not close. For roughly half the price of a DGX Spark or a 128GB Mac Studio you get the same headline 128GB to load models into, on hardware that runs everything x86 does. It isn’t the fastest of the bunch - the memory tops out around 256 GB/s, so it shares the same bandwidth ceiling as the Spark - but for the money, the amount of model you can load is frankly brilliant. If you want the most capacity per pound and you don’t fancy learning a new operating system to get it, start here.
What it’ll actually run
The Radeon 8060S iGPU borrows from that shared pool, so with the memory split set generously in the BIOS you can hand it most of the 128GB. That means a 70B model at Q4 fits with room to spare, and you can sit comfortably in the 30B to 70B range that most people top out wanting at home. Smaller models in the 8B to 14B bracket run nicely and leave you plenty of memory for context.
Speed is the caveat, same as its rivals. At around 256 GB/s the bandwidth is the bottleneck, so a large model answers at a measured pace rather than instantly - expect single figures to low double figures of tokens per second on a 70B, quicker as you drop to smaller models. There’s a 50 TOPS NPU on board too, though for local LLMs it’s not much use yet - no inference tooling had made proper use of it at the time of writing, so the graphics side does the actual work. The software picture is the reassuring part: it’s plain x86, so Windows and Linux both run, and the ROCm stack keeps improving for the AMD graphics side. You will need to nudge the memory allocation in the BIOS to get the most out of it, but that’s a one-time bit of fiddling.
Who should buy it
Buy it if value and flexibility are what you’re after. It’s the machine for the tinkerer who wants a real desktop that runs anything, loads big models, and doesn’t cost flagship money to do it. The Framework side of things means it’s more serviceable than most mini machines, too, which suits this site’s crowd nicely.
Look elsewhere if you need raw speed or CUDA. If your models fit in 24GB, a used 3090 will be faster for less, and if you’re wedded to Nvidia’s tooling this isn’t your box. But for the most model you can load without spending silly money, on an OS you already know, the Framework Desktop is hard to beat.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Biggest model that fits
gpt-oss-120b (MoE, ~60GB); BIOS memory split set high
Around 33 tokens/s and comfortably usable; up to 112GB of the 128GB can be handed to the GPU.
Dense large model
Llama 70B Q4 (~40GB)
About 5 tokens/s; it fits, but a dense 70B is slow here.
Everyday use
8B-32B Q4 (Qwen2.5 14B, Llama 8B)
Qwen2.5 14B around 23 tokens/s, Llama 8B around 21; comfortable for daily work.
Backend tip
llama.cpp with the Vulkan backend
Vulkan beat ROCm on Strix Halo in testing; Ollama defaults can run CPU-only, so force the iGPU.
What owners report
Real first-hand experience gathered from owners and the community.
- “
Single Framework Desktop 128GB tests: Llama 3.1 70B Q4 dense about 5 tokens/s, gpt-oss-120b (MoE) about 33, gpt-oss-20b about 45, Qwen2.5 14B about 23. Power draw sat around 97-140W through these runs.
- “
As of testing, the 50 TOPS Strix Halo NPU could not be used for local LLM inference (no tooling had cracked it), and the Vulkan backend outran ROCm.
- “
The 256 GB/s is the theoretical figure; real-world measured bandwidth on Strix Halo comes in lower, around 210-215 GB/s, and it is what caps token speed on large models.
2 claim(s) we couldn't fully verify
- · Memory bandwidth of 256 GB/s in practice - 256 GB/s is the theoretical maximum; community measurements land nearer 210-215 GB/s, so real-world bandwidth is lower than the quoted figure.
- · Power draw of 140W (power_w) - Reflects measured LLM-load draw (Geerling saw 97-153W); the unit ships with a larger PSU, so 140W is a load figure rather than the supply rating.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
- Benchmark Framework Desktop Mainboard (single node, 128GB)Jeff Geerling · primary source
- AMD Ryzen AI Max+ 395 product specificationsAMD · primary source
- Strix Halo (Ryzen AI Max+ 395) LLM Benchmark Resultslhl (Level1Techs forum) · forum
- Framework Desktop Review: A Solid AMD Strix HaloServeTheHome · review
Common questions
How much memory can the Framework Desktop give to models?+
Up to 128GB is shared between the CPU and the Radeon 8060S graphics. You set the split in the BIOS, and with it set generously the graphics side can take the lion's share, which is what lets a 70B model at Q4 fit with room for context.
Is the Framework Desktop fast for local LLMs?+
It loads big models well but answers at a measured pace. Memory bandwidth tops out around 256 GB/s, the same ceiling as the DGX Spark, so a 70B lands single figures to low double figures of tokens per second. Smaller models feel quicker.
Does it need CUDA to run AI models?+
No. It's an AMD machine, so it runs on ROCm and standard tools like llama.cpp and Ollama on Windows or Linux. If your workflow depends specifically on Nvidia CUDA, this isn't the box for you.
