Ryzen AI Max+ 395 (Strix Halo): 128 GB of memory for LLMs locaux
The Ryzen AI Max+ 395 (codenamed Strix Halo) is the APU AMD is using to take on local AI head-on, a space previously dominated by unified-memory Macs and 24 GB GPUs. Its key selling point: up to 128 GB of memory shared among the CPU, GPU, and NPU, with the overwhelming majority available for allocation to the model. This guide gives an honest assessment of what the Ryzen AI Max+ 395 can really do for LLMs — which 70B and MoE models become accessible, at what speed, compared with dedicated GPUs and Macs, and how to configure it under ROCm and Vulkan.
Choosing a machine? Our picks by budget → · Our spec sheet Ryzen AI Max mini-PC with 128 GB →
Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).
Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why the Ryzen AI Max+ 395 changes the game for local LLMs
Running a large model locally almost always hits the same wall: video memory. A RTX 4090 tops out at 24 GB, a RTX 5090 at 32 GB, and beyond that you need to stack multiple cards or turn to a Mac Studio. The Ryzen AI Max+ 395 attacks this wall from another angle: instead of limited dedicated VRAM, it shares a large unified memory pool between the CPU, integrated GPU, and NPU. On a model equipped with 128 GB, you can reserve more than 90 GB for the GPU during inference.
In practical terms, this opens up a category of machines that did not exist on the PC side: mini PCs and laptops at widely varying prices, which need to be checked for their ability to load a 70B model—or much larger mixtures of experts—without a dedicated graphics card and without a multi-GPU setup. It's the first credible x86 competitor to the Mac Studio for local AI with large memory capacity.
#The chip in detail
The Ryzen AI Max+ 395 is a monolithic APU that combines three compute engines on a single die, all connected to the same memory controller. For an LLM, the GPU and memory matter most, but the whole design forms a coherent unit.
- CPU — 16 Zen 5 cores
- 16 Zen 5 cores / 32 threads, useful for preprocessing, partial CPU offload, and everything other than GPU inference itself.
- GPU — Radeon 8060S (RDNA 3.5)
- 40 compute units based on the RDNA 3.5 architecture. This is the largest integrated GPU on the PC market, and it handles inference through ROCm or Vulkan.
- NPU — XDNA 2, ~50 TOPS
- A dedicated low-power AI accelerator. Relevant for some optimized workloads, but most LLM runtimes (Ollama, llama.cpp) currently target the GPU, not the NPU.
- Memory — LPDDR5X up to 128 GB
- 256-bit bus, LPDDR5X at 8000 MT/s, or approximately 256 GB/s of shared bandwidth. Unified across the CPU, GPU, and NPU.
#128 GB of unified memory, the real selling point
On a conventional GPU, VRAM and system RAM are two separate pools: when a model exceeds VRAM, it switches to RAM over PCIe, and inference collapses. On the Ryzen AI Max+ 395, there is only one physical pool. The GPU addresses unified memory directly, just like the unified memory on a Mac Apple Silicon. No copying between two address spaces, no PCIe bottleneck.
The amount actually available to the model depends on how memory is split between "system" memory and GPU memory. Depending on the firmware and OS, a fixed portion may be reserved for the GPU (the UMA setting in the BIOS) and/or the rest may be allocated dynamically by the driver (GTT under Linux). In practice, on a 128 GB machine, it is common to make 96 GB or more accessible to the GPU for inference.
- A single pool
- 128 GB shared. What the model does not use remains available to the system, and vice versa.
- Generous GPU allocation
- 90 GB and more can be allocated to model weights and the context cache, depending on the BIOS/driver configuration.
- No PCIe bottleneck
- Unlike a dedicated GPU that spills into RAM, everything here fits within the same high-speed memory space.
- Comfortable long context
- The extra memory headroom enables larger context windows where a 24 GB GPU has to sacrifice either model size or context.
#Which models fit (70B Q4, MoE)
With ~90 GB usable, the Ryzen AI Max+ 395 completely changes the list of models you can consider compared with a 16–24 GB GPU. Here are the concrete tiers, using the usual Q4 VRAM benchmarks: 7B ≈ 5 GB, 14B ≈ 9 GB, 32B ≈ 19 GB, 70B ≈ 40 GB.
- 70B in Q4_K_M (~40 GB)
- Plenty of room, with headroom for a large context. This is the flagship use case: a dense 70B model or equivalent, impossible on a single RTX 4090.
- 70B in Q8_0 (~75 GB)
- It also works, with almost no loss in quality. Unthinkable on a consumer GPU without multi-GPU.
- MoE such as Qwen 3.6 35B-A3B or GLM 4.7 Flash
- Mixture-of-experts models are ideal here: all the weights are loaded into memory (a lot), but only a few billion parameters are active per token, so speed remains good.
- Large MoE (200B+ in Q4)
- Models such as Qwen3-235B-A22B in aggressive quantization can fit in 90 GB. Throughput depends on the number of active parameters, not the total size.
- DeepSeek 671B and the rest
- Too large even in Q4—well over 100 GB for the weights. Remains out of reach without disk offload—prefer a Mac Studio with very large memory or a dedicated llama.cpp workstation.
#Real-world performance: bandwidth decides
This is the key point to understand before any purchase. Token generation for a dense model is limited by memory bandwidth, not raw compute power. The theoretical maximum throughput can be roughly calculated by dividing bandwidth by the model's size in memory.
The Ryzen AI Max+ 395 provides approximately 256 GB/s. A 70B model in Q4 (~40 GB) therefore theoretically tops out at around 6 tokens/s, while in practice generation is closer to 4 to 5 tokens/s—readable, but not fast. By contrast, a 30B-A3B MoE reads only a small portion of the weights per token and can easily reach several dozen tokens/s. Prompt preprocessing uses the GPU's 40 CUs and remains adequate.
- Dense 70B Q4 model
- ≈ 4–5 tok/s during generation. Comfortable for writing and analysis, but slow for highly interactive chat.
- 32B dense Q4 model
- Noticeably smoother (~19 GB read/token), a good everyday quality/speed compromise.
- 30B-A3B MoE
- Several dozen tok/s: speed depends on the ~3B active parameters, not the 30B total parameters.
- RTX comparison
- An RTX 4090 (~1000 GB/s) crushes the Strix Halo in raw throughput, but tops out at 24 GB and simply cannot load a 70B on its own.
#Mini-PC vs. laptop vs. Mac Studio: the honest comparison
The Ryzen AI Max+ 395 comes in several form factors. The choice comes down to mobility, sustained cooling, and price, since the chip itself remains identical. Opposite it, the benchmark to beat remains the Mac Studio and its M-Max/Ultra variants with large unified memory.
- Mini PC (e.g., Framework Desktop, GMKtec EVO-X2)
- The best sustained-performance-to-price ratio. Stable cooling for long inference runs, compact, quiet, and ideal for a 24/7 home server.
- Laptop (e.g., HP ZBook Ultra G1a, Asus ROG Flow Z13)
- The same memory capacity in a mobile form factor, but thermal throttling under sustained load and limited battery life during inference. Convenient for working anywhere, less suitable for a server.
- Mac Studio (M-Max / Ultra)
- Much higher memory bandwidth (several hundred GB/s, up to ~800 GB/s on Ultra) and a highly optimized MLX ecosystem. Faster on dense models, but significantly more expensive at equivalent memory.
- Mac mini M4 Pro
- Less maximum memory, but excellent for models up to 32B. Consider it if you don't need 70B+.
In summary: Strix Halo wins on price at high memory capacity and on the x86/Linux ecosystem; Mac Studio wins on generation speed and software maturity. For an “accessible” 70B setup without breaking the bank, the 128 GB Ryzen AI Max+ 395 mini-PC is currently the most relevant x86 option.
#ROCm and Vulkan setup on Linux
On Linux, two paths lead to the integrated GPU: ROCm (AMD's compute stack) and Vulkan (portable and often easier to get working on this newer GPU). The key step, common to both, is making enough memory accessible to the GPU.
- 01Adjust GPU memory allocation in the BIOSEnter the BIOS/UEFI and look for the UMA Frame Buffer Size setting (or « Dedicated GPU Memory »). Reserve a generous portion for the GPU. The rest will be allocated dynamically by the amdgpu driver through GTT memory.
- 02Check the GTT memory visible to the driverOn Linux, memory that can be dynamically allocated to the GPU goes through the amdgpu driver's GTT mechanism. A recent kernel (6.10+) and up-to-date amdgpu firmware maximize the share available for inference.
- 03Install a GPU-targeted stackThe simplest option is to use a build of Ollama or llama.cpp with the Vulkan backend, which detects the Radeon 8060S without heavy ROCm configuration. For ROCm, install the official package and verify that the GPU is listed correctly.
- 04Force the GPU target if necessaryBecause the integrated GPU is not always on the official list of ROCm targets, you sometimes need to specify the GFX architecture version through an environment variable so that Ollama/llama.cpp accepts it.
- 05Run a model and verify offloadingStart a model and confirm that it is actually running on the GPU (not the CPU) before drawing any conclusions about performance.
#Windows configuration
On Windows, the driver experience is more “plug and play,” but memory allocation remains the sensitive point. The logic is the same: give the GPU enough memory, then use a runtime that can take advantage of it.
- 01Install the latest AMD Adrenalin driverUse the latest driver for the Ryzen AI Max+ 395: integrated GPU LLM support and memory allocation improve with every release. Recent drivers expose a significant portion of unified memory as VRAM.
- 02Adjust dedicated graphics memoryIn the BIOS (UMA Frame Buffer Size) or through the AMD utility, increase dedicated graphics memory to make room for larger models. A setting that is too low will spill the model onto the CPU.
- 03Install LM Studio or OllamaLM Studio detects the AMD GPU and suggests a suitable backend; it’s the simplest option without a command line. Ollama for Windows also works and exposes its endpoint at http://localhost:11434.
- 04Load a GGUF and verify GPU offloadIn LM Studio, load a 70B model in Q4 and move the GPU offload slider all the way up. Make sure the layers are actually on the GPU and not the CPU.
#Troubleshooting and tips
- The model runs on the CPU, not the GPU
- Check “ollama ps” / “rocm-smi”. GPU memory allocation is often too low: increase the UMA Frame Buffer in the BIOS or check the GTT under Linux.
- ROCm does not detect the Radeon 8060S
- The integrated RDNA 3.5 GPU is not always on the official list. Export HSA_OVERRIDE_GFX_VERSION, or switch to the Vulkan backend, which detects it without overhead.
- Abnormally slow generation on a dense 70B model
- This is expected: ~4–5 tok/s is the realistic ceiling at 256 GB/s. For more speed, switch to an MoE or a smaller model—it isn't a bug.
- Unable to allocate 90 GB to the GPU
- Outdated BIOS firmware and driver/kernel limit allocation. Update the BIOS, the Adrenalin driver (Windows), or the amdgpu kernel/firmware (Linux).
- Long context that causes crashes
- The KV cache is added to the weights in memory. Reduce num_ctx or the model size if you approach the ~90 GB limit.
#Machines featuring Strix Halo (August 2026)
The chip's success is reflected in its catalog: more than a dozen machines already use it, from aggressive mini PCs to branded workstations, portable consoles, and modular boxes. The indicative prices below are those observed in August 2026 for 128 GB configurations—they change quickly.
| Machine | Format | Special feature | Indicative price |
|---|---|---|---|
| Bosgame M5 | Mini-PC | The least expensive in the segment | ~1 700 $ |
| GMKtec EVO-X2 | Mini-PC | 4 8K outputs | highly variable price, verify it |
| Beelink GTR9 Pro | Mini-PC | Dual 10 GbE | ~2 000 $ |
| Framework Desktop | Modular case | Repairable, beloved by developers | highly variable price, verify it |
| Corsair AI Workstation 300 | SFF | Established brand, solid warranty | 2 000-3 400 $ |
| Minisforum MS-S1 MAX | Mini-PC | PCIe x16 slot: the only upgradeable one | ~2 500 $ |
| HP Z2 Mini G1a | 2.7 L workstation | ISV-certified, pro channel (PRO 395) | 2 400-3 500 $ |
| Geekom A9 Mega | Mini-PC | 2 TB of original SSD storage | ~3 200 $ |
| Peladn YO1 | 2.4 L mini PC | Sold as “deploy a 70B” | ~1 850 $ |
| Acemagic M1A Pro+ | Mini-PC | Up to 24 TB of storage | n.c. |
| Aoostar NEX395 | Mini-PC | 140 W TDP (the most powerful) | China, EU coming soon |
| GPD Win 5 | Portable console | A 70B in your hands | announced |
| Minix ER939-AI | Mini-PC | Dual 10 GbE variant | n.c. |
Our buying advice fits in one line: with identical specs (the chip is the same everywhere), use price, warranty, and cooling as the tie-breakers. Tests of each machine are gradually being added to our Reviews & Tests section.
#Frequently asked questions
Can the 128 GB of memory really be used by an LLM?+
What is the largest model that can run on Strix Halo?+
Can Strix Halo replace a RTX 5090?+
Which Strix Halo machine should you choose?+
Windows or Linux for local AI on Strix Halo?+
#Go further
The Ryzen AI Max+ 395 is part of a broader look at hardware and models for local AI. These site guides build on this one:
- Choose your GPU for local AI
- The buying guide that puts Strix Halo in perspective against RTX and Mac, based on your budget and target models.
- Ollama with AMD GPUs (ROCm)
- The detailed ROCm configuration, useful for getting the most out of the Radeon 8060S once you have the platform in hand.
- Choose your quantization (Q4, Q5, Q8, FP16)
- For choosing between fast Q4 and near-lossless Q8 when memory is plentiful—exactly the case with 128 GB.
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.