AMD Ryzen AI Max / Strix Halo Deep Dive: The Aggressive Bet on Unified Memory
> Chip: AMD Ryzen AI Max+ 395 (Strix Halo) > Architecture: 4nm SoC, Zen 5 + RDNA 3.5 + XDNA 2 > Memory: Up to 128GB LPDDR5x-8000, 256-bit bus > Launch: Q2 2026 > Next-gen: Gorgon Halo (192GB unified memory, 160GB VRAM)
Key specifications
| Component | Spec | Notes | |---|---|---| | CPU | 16x Zen 5 cores / 32 threads | Up to 5.1 GHz, 80MB L2+L3 | | GPU | 40x RDNA 3.5 CU | Radeon 8060S, up to 2900 MHz | | NPU | XDNA 2 | 50 TOPS rated (see below) | | Memory | 256-bit LPDDR5x-8000 | Unified, up to 128GB | | Bandwidth | ~256 GB/s | Shared by CPU + GPU + NPU | | TDP | 45-120W | Up to 125W, fits in laptops | | Process | TSMC 4nm | — |
Key points
- Unified memory architecture: For the first time in the x86 ecosystem, CPU and GPU share one pool of LPDDR5x-8000. Up to 96GB can be allocated as VRAM via BIOS. There is no PCIe copy overhead, and pointers can be passed directly.
- Why it matters for AI: An RTX 5080 has only 16GB VRAM — full DeepSeek R1 (671B total, 235B active) needs 150GB+. Strix Halo's 96GB VRAM can hold 70B dense models at Q4 quantization (~40-45GB) and parts of MoE models. The pitch is not "runs fast" but "runs at all" on consumer hardware.
- The "3x faster than RTX 5080" claim is a capacity result, not a speed result. The 5080 spills to system memory over PCIe and collapses; Strix Halo runs slowly but completes. On models that fit in 16GB (7B/14B), the 5080 wins on GDDR7 bandwidth and CUDA throughput.
- MoE (DeepSeek R1): ~50 t/s — smooth, beyond reading speed. Only ~37B parameters activate per token (~74GB at FP16, quantized to fit), and the 32MB MALL cache intercepts most memory traffic.
- Dense (Llama 3.3 70B): 5-6 t/s — noticeably stuttery. Every token reads all weights; 256 GB/s ÷ ~40GB (Q4) ≈ 6.4 t/s theoretical ceiling, essentially reached.
- AMD official specs: https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/ryzen-ai-max-395.html
- Ultrabook Review: https://www.ultrabookreview.com/70442-amd-strix-halo-laptops/
- VideoCardz (Gorgon Halo): https://videocardz.com/newz/amd-confirms-ryzen-ai-max-400-gorgon-halo-will-support-up-to-192gb-memory-and-160gb-vram
- Toolhalla (local LLM guide): https://toolhalla.ai/blog/amd-strix-halo-local-llm-guide-2026
- GitHub (ROCm benchmarks): https://github.com/nabe2030/faster-whisper-rocm-strix-halo
The pricing trap
| Configuration | Memory | Price | Runs 235B model? | |---|---|---|---| | GMKtec EVO-X2 (base) | 64GB | ~$1,499 | No | | GMKtec EVO-X2 (high) | 128GB | ~$2,199-2,299 | Yes | | Framework Desktop | From 32GB | From $1,099 | No (32GB) | | AMD Ryzen AI Halo dev PC | 128GB | $3,999 | Yes | | Asus ROG Flow Z13 | From 32GB | $2,199 | No (32GB) |
The machine shown in the demo is the 128GB model (~$2,200), not the $1,499 64GB base unit.
Inference performance: MoE vs dense polarization
NPU: hardware ready, software lagging
The 50 TOPS XDNA 2 NPU manages only 4.4 t/s on Llama 3.2 1B. Analysis shows ~75% of time is driver/scheduling overhead; only ~25% is actual tensor compute. The ROCm + XDNA software stack remains far behind CUDA. The RDNA 3.5 GPU is the practical workhorse for LLMs today.
Versus Apple M4 Max
| Dimension | Ryzen AI Max+ 395 | Apple M4 Max | Winner | |---|---|---|---| | Unified memory bandwidth | 256 GB/s | 546 GB/s | Apple | | 70B dense inference | 5-6 t/s | 15-25 t/s | Apple | | MoE inference | ~50 t/s | Similar/slightly better | Roughly even | | 128GB price | ~$1,999-3,299 | $3,699 | AMD | | OS | Linux + Windows | macOS | AMD | | Gaming (1080p Ultra) | 75-85 fps | Similar | Even | | Single-core efficiency | Behind | Ahead | Apple |
Apple's 512-bit memory bus is the key gap. AMD's counterpoints: lower price, Linux support with seamless local-to-cloud Docker migration, and x86 ecosystem consistency.
Gaming performance
Roughly laptop RTX 4070 / desktop RTX 3060 Ti class at 1080p Ultra: Cyberpunk 2077 (RT) 75.6 fps, Baldur's Gate 3 85.3 fps, GTA V 83.5 fps, Horizon Zero Dawn ~70-80 fps.
The payback math is flawed
The viral "9-month payback vs $5,280/year cloud subscription" comparison wrongly substitutes a cloud subscription price for local-machine replacement cost. A more honest estimate: if you spend $400/month on AI services and migrate ~$200/month of usage locally, payback is 11-18 months — frontier-model tasks will keep you on the cloud. Local inference complements rather than replaces cloud services.
Next-gen: Gorgon Halo
The confirmed Ryzen AI Max 400 series (Gorgon Halo) raises unified memory to 192GB and VRAM allocation to 160GB, with NPU up to 55 TOPS and GPU clocks to 3000 MHz; CPU/GPU core counts stay the same. It could run larger dense models (~120B class) and enable multi-model concurrency, but pricing is unknown and likely above $3,000.
Verdict: who should buy
Good fit: local AI developers running 70B+ models; privacy-sensitive users (medical/legal/financial); MoE model enthusiasts; anyone needing x86 + Linux with cloud-portable containers; mobile workstation users needing CPU+GPU+NPU with 128GB.
Bad fit: dense-model speed requirements (buy a Mac Studio M4 Max or RTX 4090/5090 instead); pure value seekers ($1,499 units can't run large models); gamers; NPU-dependent workloads; students on a budget.
The unified-memory bet paid off: the chip is excellent and brings enterprise-scale AI inference to consumer hardware. But the NPU software stack, limited bandwidth for dense models, DRAM-driven pricing, and next-gen competition remain real weaknesses.