AI PC · mini-pc
GMKtec EVO-X2 (Ryzen AI Max+ 395)
128GB of Strix Halo on Windows - though a 2026 memory shortage has knocked it off its value perch.
- Memory
- 128 GB
- Bandwidth
- 256 GB/s
- AI compute
- 50 TOPS
- £ / GB
- £25
What it runs
With 128 GB of memory you can load, at 4-bit, a model up to roughly
~252B params
Verdict
A 128GB Strix Halo box that loads a 70B - once the value pick, but a 2026 memory shortage has pushed it above the Beelink and Framework it used to undercut.
BEST FOR
- + big models on Windows
- + buyers who catch it on one of GMKtec's frequent sales
NOT FOR
- - fast tokens on dense models
- - anyone who needs CUDA
- - value hunters at the current shortage price
The GMKtec EVO-X2 takes the Ryzen AI Max+ 395 and its 128GB of shared memory and wraps it in a mini-PC that, for most of its life, undercut nearly everything else with the same chip. That changed in 2026: an LPDDR5X memory shortage spiked its price to around $3,400-3,650, above the Beelink and Framework it used to beat. The capacity is still superb; the value crown has slipped, and it now depends on catching a GMKtec sale.
The take
This used to be the value pick of the 128GB boxes, plain and simple - same Strix Halo chip as the Beelink, same 128GB you can hand mostly to the graphics side, and it routinely sold for around half what the premium rivals asked. The 2026 memory shortage flipped that: the EVO-X2’s 128GB SKU jumped to around $3,400-3,650 while the Beelink held near $1,999, so today it’s the dearer of the two, not the cheaper. It still won’t win a speed contest - the memory tops out near 256 GB/s in theory and lower in practice, so a dense 70B answers at a walk - and the capacity remains superb. But buy it now for capacity, and only for value if you catch one of GMKtec’s regular sales that drops it back toward its old price.
What it’ll actually run
The Radeon 8060S draws on the shared pool, so with the memory split set generously you can put most of the 128GB toward models. That means a 70B at Q4 fits with room for context, and big mixture-of-experts models are where it shines. GMKtec’s own figures on the 128GB unit have gpt-oss-120b around 19 tokens a second, a Qwen3 30B MoE near 55, and even a 235B MoE ticking over about 11 - impressive for the money. Dense models are the slow lane: a DeepSeek R1 32B lands under 10 tokens a second, so temper expectations there.
The bottleneck is bandwidth, same as every Strix Halo box. Around 256 GB/s on paper, and real-world measurements come in lower, so dense large models feel measured rather than quick. There’s a 50 TOPS NPU on board too, but no local LLM tooling makes proper use of it yet, so the GPU does the work. Software is the reassuring part: it ships with Windows 11, runs Linux happily, and the AMD stack keeps improving. You’ll want to set the memory split in the BIOS to feed the GPU, but that’s a one-time job.
Who should buy it
Buy it if capacity is the point and you catch it at a sensible price. It’s still a fine machine for loading big models at home on Windows or Linux, and at a sale price it’s back to being the value champion. At the current shortage price, though, look hard at the Beelink GTR9 Pro first, which holds the same 128GB for less right now.
Look elsewhere if you need speed or CUDA. If your models fit in 24GB a used 3090 runs them faster for less, and if you’re tied to Nvidia’s tooling none of it lives here. The EVO-X2 is still the most model you can load in a tidy little Windows box - just no longer automatically the cheapest way to do it.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Biggest model that fits
gpt-oss-120b MoE (~60GB); BIOS memory split set high
GMKtec's 128GB figure is about 19 tokens/s; up to 96-112GB of the pool can go to the GPU.
Everyday sweet spot
Qwen3 30B MoE, 8B-32B Q4
Qwen3 30B around 55 tokens/s, gpt-oss-20b around 57; comfortable for daily work.
Dense large model
DeepSeek R1 32B / Llama 70B Q4
Dense 32B under 10 tokens/s; a dense 70B fits but drags. MoE models are far happier here.
Backend tip
llama.cpp with the Vulkan backend on Windows, or Lemonade on Linux
Vulkan tends to beat ROCm on Strix Halo; set the BIOS memory split first.
What owners report
Real first-hand experience gathered from owners and the community.
- “
GMKtec's published 128GB benchmarks: gpt-oss-120b about 19.25 tokens/s, gpt-oss-20b about 57, Qwen3 30B about 55, DeepSeek R1 32B about 9.3, and a 235B MoE about 11. MoE models fly, dense models drag.
- “
The 256 GB/s is the theoretical ceiling for Strix Halo; real-world measured bandwidth lands lower, around 210-215 GB/s, and that is what caps token speed on large dense models.
- “
The 50 TOPS NPU can't yet be used for local LLM inference; the Radeon 8060S iGPU does the work, and Vulkan often beats ROCm on this chip.
3 claim(s) we couldn't fully verify
- · memory_bandwidth_gbs 256 - GMKtec doesn't publish a bandwidth figure; 256 GB/s is the Strix Halo platform theoretical (LPDDR5X-8000, 256-bit), and real-world measurements land nearer 210-215 GB/s.
- · power_w 140 - 140W is the vendor's peak figure; sustained is 120W, and the bundled charger is rated about 230W.
- · Price around GBP 3,200 / USD 3,500 (Sept 2026) - Grounded Sept 2026 sweep: the 128GB EVO-X2 spiked to ~USD 3,400-3,650 (from ~1,999) in the 2026 LPDDR5X memory shortage, now ABOVE the Beelink GTR9 Pro (~1,999) - reversing its value advantage. GMKtec runs frequent sales, so it's volatile; the value thesis is reframed around that in the body/verdict/FAQ.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
- GMKtec EVO-X2 product page (vendor spec + 128GB benchmarks)GMKtec · primary source
- AMD Ryzen AI Max+ 395 product specificationsAMD · primary source
- Strix Halo (Ryzen AI Max+ 395) LLM Benchmark Resultslhl (Level1Techs forum) · forum
Common questions
How much memory can the EVO-X2 give to models?+
Up to 128GB is shared between the CPU and the Radeon 8060S graphics. You set the split in the BIOS, and with it set generously the graphics side can take most of the pool, which is what lets a 70B at Q4 fit with room for context.
Is the GMKtec EVO-X2 fast for local LLMs?+
It depends on the model. Mixture-of-experts models are quick - GMKtec's own figures have gpt-oss-120b near 19 tokens a second - but dense models are the slow lane, with a DeepSeek R1 32B under 10. Memory tops out near 256 GB/s in theory and lower in practice, so that's the ceiling.
Does it need CUDA to run AI models?+
No. It's an AMD machine, so it runs on Vulkan or ROCm with standard tools like llama.cpp and Ollama on Windows or Linux. If your workflow depends specifically on Nvidia CUDA, this isn't the box for you.
Is the EVO-X2 still cheaper than the Beelink GTR9 Pro?+
Not right now. It used to undercut the Beelink by a wide margin on the same chip and 128GB, but the 2026 LPDDR5X memory shortage spiked the EVO-X2 to around $3,400-3,650 while the Beelink held near $1,999. So the value crown has passed to the Beelink; the EVO-X2's case now rests on GMKtec's frequent sales bringing it back down.
