AMD Ryzen AI Max+ 395 (128GB unified, Strix Halo)

Compare AMD Ryzen AI Max+ 395 (128GB unified, Strix Halo) against another machine in the Arena

Class
APU
Memory
128 GB LPDDR5X-8000
Runs (Q4_K_M)
Bandwidth
256 GB/s
TDP
120 W
Released
2025-03-17
Price (US)
$1999.99 new as of 2026-08
Sourcing and disambiguation notes

40 CU RDNA 3.5 iGPU (Radeon 8060S) + 16-core Zen 5 CPU + XDNA2 NPU (50 TOPS int8, AMD-quoted, not independently verified against an LLM workload). AMD-quoted memory bandwidth is 256 GB/s theoretical; community measurement puts real achieved bandwidth at ~215 GB/s. The GPU cannot address all 128GB: the BIOS/UMA setting caps dedicated VRAM at 96GB per node, and the community benchmarks below allocate GPU memory via GTT/kernel parameters instead, reaching ~110-120GB usable. ROCm for gfx1151 is available via TheRock nightly builds and, since ROCm 6.4.4/7.0, official releases; Vulkan is the lower-friction default and is generally faster for short-context decode, while ROCm pulls ahead at long context with Flash Attention. ROCm 7.0.1 showed a severe prefill regression against 6.4.4 on this chip, so the backend VERSION, not just the backend, materially changes the results below. The recorded price is GMKtec's EVO-X2 sale price for this 128GB/2TB configuration; GMKtec lists it at $2,199.99 and other retailers have shown $1,700-$2,300. The Framework Desktop, the machine behind the original $1,999 launch figure, is reported to have risen to roughly $3,449, which could not be confirmed against Framework's own order page.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 128 GB; models that fit are listed largest first.