AI PC · mini-pc
GMKtec EVO-X2 (Ryzen AI Max+ 395)
128GB of Strix Halo for the least money: the value way into big local models on Windows.
- Memory
- 128 GB
- Bandwidth
- 256 GB/s
- AI compute
- 50 TOPS
- £ / GB
- £13
What it runs
With 128 GB of memory you can load, at 4-bit, a model up to roughly
~252B params
Verdict
The cheapest way to 128GB of unified memory: a Strix Halo box that loads a 70B for Framework money or less.
BEST FOR
- + most memory per pound
- + big models on Windows
NOT FOR
- - fast tokens on dense models
- - anyone who needs CUDA
The GMKtec EVO-X2 takes the Ryzen AI Max+ 395 and its 128GB of shared memory and wraps it in a mini-PC that undercuts nearly everything else with the same chip. If you want the most model you can load for the least outlay, and you’re happy on Windows or Linux, this is where the value lives.
The take
This is the value pick of the 128GB boxes, plain and simple. Same Strix Halo chip as the pricier Beelink, same 128GB you can hand mostly to the graphics side, but it routinely sells for around half what the premium rivals ask. It won’t win a speed contest - the memory tops out near 256 GB/s in theory and lower in practice, so a dense 70B answers at a walk - but the amount of model you can load per pound is the best going. Buy it for capacity and value, not for pace.
What it’ll actually run
The Radeon 8060S draws on the shared pool, so with the memory split set generously you can put most of the 128GB toward models. That means a 70B at Q4 fits with room for context, and big mixture-of-experts models are where it shines. GMKtec’s own figures on the 128GB unit have gpt-oss-120b around 19 tokens a second, a Qwen3 30B MoE near 55, and even a 235B MoE ticking over about 11 - impressive for the money. Dense models are the slow lane: a DeepSeek R1 32B lands under 10 tokens a second, so temper expectations there.
The bottleneck is bandwidth, same as every Strix Halo box. Around 256 GB/s on paper, and real-world measurements come in lower, so dense large models feel measured rather than quick. There’s a 50 TOPS NPU on board too, but no local LLM tooling makes proper use of it yet, so the GPU does the work. Software is the reassuring part: it ships with Windows 11, runs Linux happily, and the AMD stack keeps improving. You’ll want to set the memory split in the BIOS to feed the GPU, but that’s a one-time job.
Who should buy it
Buy it if value and capacity are the point. It’s the machine for someone who wants to load big models at home, doesn’t want to spend flagship money, and is fine on Windows or Linux. For a first serious local-AI box that can actually hold a 70B, it’s hard to argue with the price.
Look elsewhere if you need speed or CUDA. If your models fit in 24GB a used 3090 runs them faster for less, and if you’re tied to Nvidia’s tooling none of it lives here. But for the most memory per pound in a tidy little Windows machine, the EVO-X2 is the value answer.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Biggest model that fits
gpt-oss-120b MoE (~60GB); BIOS memory split set high
GMKtec's 128GB figure is about 19 tokens/s; up to 96-112GB of the pool can go to the GPU.
Everyday sweet spot
Qwen3 30B MoE, 8B-32B Q4
Qwen3 30B around 55 tokens/s, gpt-oss-20b around 57; comfortable for daily work.
Dense large model
DeepSeek R1 32B / Llama 70B Q4
Dense 32B under 10 tokens/s; a dense 70B fits but drags. MoE models are far happier here.
Backend tip
llama.cpp with the Vulkan backend on Windows, or Lemonade on Linux
Vulkan tends to beat ROCm on Strix Halo; set the BIOS memory split first.
What owners report
Real first-hand experience gathered from owners and the community.
- “
GMKtec's published 128GB benchmarks: gpt-oss-120b about 19.25 tokens/s, gpt-oss-20b about 57, Qwen3 30B about 55, DeepSeek R1 32B about 9.3, and a 235B MoE about 11. MoE models fly, dense models drag.
- “
The 256 GB/s is the theoretical ceiling for Strix Halo; real-world measured bandwidth lands lower, around 210-215 GB/s, and that is what caps token speed on large dense models.
- “
The 50 TOPS NPU can't yet be used for local LLM inference; the Radeon 8060S iGPU does the work, and Vulkan often beats ROCm on this chip.
3 claim(s) we couldn't fully verify
- · memory_bandwidth_gbs 256 - GMKtec doesn't publish a bandwidth figure; 256 GB/s is the Strix Halo platform theoretical (LPDDR5X-8000, 256-bit), and real-world measurements land nearer 210-215 GB/s.
- · power_w 140 - 140W is the vendor's peak figure; sustained is 120W, and the bundled charger is rated about 230W.
- · Price around GBP 1,699 / USD 1,799 - Indicative; GMKtec runs frequent sales and the 128GB variant's price moves. List has been up to USD 1,999.99.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
- GMKtec EVO-X2 product page (vendor spec + 128GB benchmarks)GMKtec · primary source
- AMD Ryzen AI Max+ 395 product specificationsAMD · primary source
- Strix Halo (Ryzen AI Max+ 395) LLM Benchmark Resultslhl (Level1Techs forum) · forum
Common questions
How much memory can the EVO-X2 give to models?+
Up to 128GB is shared between the CPU and the Radeon 8060S graphics. You set the split in the BIOS, and with it set generously the graphics side can take most of the pool, which is what lets a 70B at Q4 fit with room for context.
Is the GMKtec EVO-X2 fast for local LLMs?+
It depends on the model. Mixture-of-experts models are quick - GMKtec's own figures have gpt-oss-120b near 19 tokens a second - but dense models are the slow lane, with a DeepSeek R1 32B under 10. Memory tops out near 256 GB/s in theory and lower in practice, so that's the ceiling.
Does it need CUDA to run AI models?+
No. It's an AMD machine, so it runs on Vulkan or ROCm with standard tools like llama.cpp and Ollama on Windows or Linux. If your workflow depends specifically on Nvidia CUDA, this isn't the box for you.
How is it cheaper than the Beelink GTR9 Pro?+
Same chip, same 128GB, but GMKtec prices the EVO-X2 far lower on its Western store - often around half what the Beelink asks for near-identical hardware. You give up the Beelink's dual 10GbE and premium build, not the model capacity.
