Intermediate 14 minAPU

Ryzen AI Max+ 395 (Strix Halo): 128 GB of memory for LLMs locaux

The Ryzen AI Max+ 395 (codenamed Strix Halo) is the APU AMD is using to take on local AI head-on, a space previously dominated by unified-memory Macs and 24 GB GPUs. Its key selling point: up to 128 GB of memory shared among the CPU, GPU, and NPU, with the overwhelming majority available for allocation to the model. This guide gives an honest assessment of what the Ryzen AI Max+ 395 can really do for LLMs — which 70B and MoE models become accessible, at what speed, compared with dedicated GPUs and Macs, and how to configure it under ROCm and Vulkan.

Choosing a machine? Our picks by budget → · Our spec sheet Ryzen AI Max mini-PC with 128 GB →

By Samir K.·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the Ryzen AI Max+ 395 changes the game for local LLMs

Running a large model locally almost always hits the same wall: video memory. A RTX 4090 tops out at 24 GB, a RTX 5090 at 32 GB, and beyond that you need to stack multiple cards or turn to a Mac Studio. The Ryzen AI Max+ 395 attacks this wall from another angle: instead of limited dedicated VRAM, it shares a large unified memory pool between the CPU, integrated GPU, and NPU. On a model equipped with 128 GB, you can reserve more than 90 GB for the GPU during inference.

In practical terms, this opens up a category of machines that did not exist on the PC side: mini PCs and laptops at widely varying prices, which need to be checked for their ability to load a 70B model—or much larger mixtures of experts—without a dedicated graphics card and without a multi-GPU setup. It's the first credible x86 competitor to the Mac Studio for local AI with large memory capacity.

i
Strix Halo, Ryzen AI Max, Radeon 8060S: what are we talking about?
“Strix Halo” is the generation’s codename. “Ryzen AI Max+ 395” is the high-end model in the Ryzen AI Max series. Its integrated GPU is called Radeon 8060S. All three refer to facets of the same chip—you will encounter all three names depending on the product specifications.

#The chip in detail

The Ryzen AI Max+ 395 is a monolithic APU that combines three compute engines on a single die, all connected to the same memory controller. For an LLM, the GPU and memory matter most, but the whole design forms a coherent unit.

CPU — 16 Zen 5 cores
16 Zen 5 cores / 32 threads, useful for preprocessing, partial CPU offload, and everything other than GPU inference itself.
GPU — Radeon 8060S (RDNA 3.5)
40 compute units based on the RDNA 3.5 architecture. This is the largest integrated GPU on the PC market, and it handles inference through ROCm or Vulkan.
NPU — XDNA 2, ~50 TOPS
A dedicated low-power AI accelerator. Relevant for some optimized workloads, but most LLM runtimes (Ollama, llama.cpp) currently target the GPU, not the NPU.
Memory — LPDDR5X up to 128 GB
256-bit bus, LPDDR5X at 8000 MT/s, or approximately 256 GB/s of shared bandwidth. Unified across the CPU, GPU, and NPU.
!
The memory is soldered on; choose it when purchasing
The Strix Halo's LPDDR5X is soldered to the board: there are no SO-DIMM slots, and no expansion is possible later. If you're targeting 70B+ models, choose the 128 GB variant from the start. The 32 or 64 GB versions undermine the platform's main advantage for large LLMs.
i
Why this chip is historic
For ten years, running a large model at home meant stacking overpriced GPUs to scrape together enough VRAM. Strix Halo changes the game: AMD combines 16 Zen 5 cores, a RTX 4060-4070-class GPU, and a 50 TOPS NPU around a single 128 GB pool—and suddenly a quantized 70B model fits entirely where a RTX 5090 costing approximately €5,700 in late September 2026 tops out at 32 GB. It’s the first consumer chip designed for local inference, rather than adapted for it.

#128 GB of unified memory, the real selling point

On a conventional GPU, VRAM and system RAM are two separate pools: when a model exceeds VRAM, it switches to RAM over PCIe, and inference collapses. On the Ryzen AI Max+ 395, there is only one physical pool. The GPU addresses unified memory directly, just like the unified memory on a Mac Apple Silicon. No copying between two address spaces, no PCIe bottleneck.

The amount actually available to the model depends on how memory is split between "system" memory and GPU memory. Depending on the firmware and OS, a fixed portion may be reserved for the GPU (the UMA setting in the BIOS) and/or the rest may be allocated dynamically by the driver (GTT under Linux). In practice, on a 128 GB machine, it is common to make 96 GB or more accessible to the GPU for inference.

A single pool
128 GB shared. What the model does not use remains available to the system, and vice versa.
Generous GPU allocation
90 GB and more can be allocated to model weights and the context cache, depending on the BIOS/driver configuration.
No PCIe bottleneck
Unlike a dedicated GPU that spills into RAM, everything here fits within the same high-speed memory space.
Comfortable long context
The extra memory headroom enables larger context windows where a 24 GB GPU has to sacrifice either model size or context.

#Which models fit (70B Q4, MoE)

With ~90 GB usable, the Ryzen AI Max+ 395 completely changes the list of models you can consider compared with a 16–24 GB GPU. Here are the concrete tiers, using the usual Q4 VRAM benchmarks: 7B ≈ 5 GB, 14B ≈ 9 GB, 32B ≈ 19 GB, 70B ≈ 40 GB.

70B in Q4_K_M (~40 GB)
Plenty of room, with headroom for a large context. This is the flagship use case: a dense 70B model or equivalent, impossible on a single RTX 4090.
70B in Q8_0 (~75 GB)
It also works, with almost no loss in quality. Unthinkable on a consumer GPU without multi-GPU.
MoE such as Qwen 3.6 35B-A3B or GLM 4.7 Flash
Mixture-of-experts models are ideal here: all the weights are loaded into memory (a lot), but only a few billion parameters are active per token, so speed remains good.
Large MoE (200B+ in Q4)
Models such as Qwen3-235B-A22B in aggressive quantization can fit in 90 GB. Throughput depends on the number of active parameters, not the total size.
DeepSeek 671B and the rest
Too large even in Q4—well over 100 GB for the weights. Remains out of reach without disk offload—prefer a Mac Studio with very large memory or a dedicated llama.cpp workstation.
→
MoE models are made for this machine
A dense 70B model reads 40 GB of weights for every token, so bandwidth becomes the bottleneck. An MoE with the same footprint but only 3B active parameters reads just a fraction of the memory per token, so it runs much faster. On Strix Halo, favor MoE architectures when speed matters.

#Real-world performance: bandwidth decides

This is the key point to understand before any purchase. Token generation for a dense model is limited by memory bandwidth, not raw compute power. The theoretical maximum throughput can be roughly calculated by dividing bandwidth by the model's size in memory.

The Ryzen AI Max+ 395 provides approximately 256 GB/s. A 70B model in Q4 (~40 GB) therefore theoretically tops out at around 6 tokens/s, while in practice generation is closer to 4 to 5 tokens/s—readable, but not fast. By contrast, a 30B-A3B MoE reads only a small portion of the weights per token and can easily reach several dozen tokens/s. Prompt preprocessing uses the GPU's 40 CUs and remains adequate.

Dense 70B Q4 model
≈ 4–5 tok/s during generation. Comfortable for writing and analysis, but slow for highly interactive chat.
32B dense Q4 model
Noticeably smoother (~19 GB read/token), a good everyday quality/speed compromise.
30B-A3B MoE
Several dozen tok/s: speed depends on the ~3B active parameters, not the 30B total parameters.
RTX comparison
An RTX 4090 (~1000 GB/s) crushes the Strix Halo in raw throughput, but tops out at 24 GB and simply cannot load a 70B on its own.
i
Capacity vs. speed: the right tradeoff
The Ryzen AI Max+ 395 buys memory capacity, not speed. If your need is “run a large model that no 24 GB card can load,” it's excellent. If your need is “maximum tokens/s on a 7B–14B,” a dedicated GPU remains faster and cheaper.

#Mini-PC vs. laptop vs. Mac Studio: the honest comparison

The Ryzen AI Max+ 395 comes in several form factors. The choice comes down to mobility, sustained cooling, and price, since the chip itself remains identical. Opposite it, the benchmark to beat remains the Mac Studio and its M-Max/Ultra variants with large unified memory.

Mini PC (e.g., Framework Desktop, GMKtec EVO-X2)
The best sustained-performance-to-price ratio. Stable cooling for long inference runs, compact, quiet, and ideal for a 24/7 home server.
Laptop (e.g., HP ZBook Ultra G1a, Asus ROG Flow Z13)
The same memory capacity in a mobile form factor, but thermal throttling under sustained load and limited battery life during inference. Convenient for working anywhere, less suitable for a server.
Mac Studio (M-Max / Ultra)
Much higher memory bandwidth (several hundred GB/s, up to ~800 GB/s on Ultra) and a highly optimized MLX ecosystem. Faster on dense models, but significantly more expensive at equivalent memory.
Mac mini M4 Pro
Less maximum memory, but excellent for models up to 32B. Consider it if you don't need 70B+.

In summary: Strix Halo wins on price at high memory capacity and on the x86/Linux ecosystem; Mac Studio wins on generation speed and software maturity. For an “accessible” 70B setup without breaking the bank, the 128 GB Ryzen AI Max+ 395 mini-PC is currently the most relevant x86 option.

#ROCm and Vulkan setup on Linux

On Linux, two paths lead to the integrated GPU: ROCm (AMD's compute stack) and Vulkan (portable and often easier to get working on this newer GPU). The key step, common to both, is making enough memory accessible to the GPU.

  1. 01
    Adjust GPU memory allocation in the BIOS
    Enter the BIOS/UEFI and look for the UMA Frame Buffer Size setting (or « Dedicated GPU Memory »). Reserve a generous portion for the GPU. The rest will be allocated dynamically by the amdgpu driver through GTT memory.
  2. 02
    Check the GTT memory visible to the driver
    On Linux, memory that can be dynamically allocated to the GPU goes through the amdgpu driver's GTT mechanism. A recent kernel (6.10+) and up-to-date amdgpu firmware maximize the share available for inference.
  3. 03
    Install a GPU-targeted stack
    The simplest option is to use a build of Ollama or llama.cpp with the Vulkan backend, which detects the Radeon 8060S without heavy ROCm configuration. For ROCm, install the official package and verify that the GPU is listed correctly.
  4. 04
    Force the GPU target if necessary
    Because the integrated GPU is not always on the official list of ROCm targets, you sometimes need to specify the GFX architecture version through an environment variable so that Ollama/llama.cpp accepts it.
  5. 05
    Run a model and verify offloading
    Start a model and confirm that it is actually running on the GPU (not the CPU) before drawing any conclusions about performance.
Terminal — check the GPU and launch a model
# Vérifier que le GPU AMD est détecté par ROCm
rocminfo | grep -i "gfx"

# Voir l'occupation mémoire du GPU en temps réel
rocm-smi

# Lancer un gros MoE 2026 avec Ollama
ollama run qwen3.6:35b

# Confirmer que le modèle est sur GPU (et pas CPU)
ollama ps
Terminal — force the GFX target for ROCm (if not detected)
# Le Radeon 8060S (RDNA 3.5) n'est pas toujours reconnu par défaut.
# Indiquez la version GFX pour qu'Ollama accepte le GPU.
export HSA_OVERRIDE_GFX_VERSION=11.0.0
ollama serve
→
Vulkan first, ROCm next
On this very recent platform, llama.cpp/Ollama's Vulkan backend is often the shortest path to a working GPU: fewer dependencies and automatic detection. Keep ROCm for fine-tuning performance once everything works. The site's Vulkan guide covers the detailed compilation steps.

#Windows configuration

On Windows, the driver experience is more “plug and play,” but memory allocation remains the sensitive point. The logic is the same: give the GPU enough memory, then use a runtime that can take advantage of it.

  1. 01
    Install the latest AMD Adrenalin driver
    Use the latest driver for the Ryzen AI Max+ 395: integrated GPU LLM support and memory allocation improve with every release. Recent drivers expose a significant portion of unified memory as VRAM.
  2. 02
    Adjust dedicated graphics memory
    In the BIOS (UMA Frame Buffer Size) or through the AMD utility, increase dedicated graphics memory to make room for larger models. A setting that is too low will spill the model onto the CPU.
  3. 03
    Install LM Studio or Ollama
    LM Studio detects the AMD GPU and suggests a suitable backend; it’s the simplest option without a command line. Ollama for Windows also works and exposes its endpoint at http://localhost:11434.
  4. 04
    Load a GGUF and verify GPU offload
    In LM Studio, load a 70B model in Q4 and move the GPU offload slider all the way up. Make sure the layers are actually on the GPU and not the CPU.
PowerShell — Ollama on Windows
# Vérifier la version d'Ollama
ollama --version

# Lancer un gros MoE 2026 ; l'endpoint local reste sur :11434
ollama run qwen3.6:35b

# Vérifier l'exécution sur GPU
ollama ps
i
Recommended stack
The most comfortable combination is still Ollama (the inference daemon, an OpenAI-compatible endpoint on :11434) plus an interface such as Open WebUI or LM Studio. Nothing here is specific to Strix Halo: once the GPU is being used, the rest of the workflow is identical to that of any other machine.

#Troubleshooting and tips

The model runs on the CPU, not the GPU
Check “ollama ps” / “rocm-smi”. GPU memory allocation is often too low: increase the UMA Frame Buffer in the BIOS or check the GTT under Linux.
ROCm does not detect the Radeon 8060S
The integrated RDNA 3.5 GPU is not always on the official list. Export HSA_OVERRIDE_GFX_VERSION, or switch to the Vulkan backend, which detects it without overhead.
Abnormally slow generation on a dense 70B model
This is expected: ~4–5 tok/s is the realistic ceiling at 256 GB/s. For more speed, switch to an MoE or a smaller model—it isn't a bug.
Unable to allocate 90 GB to the GPU
Outdated BIOS firmware and driver/kernel limit allocation. Update the BIOS, the Adrenalin driver (Windows), or the amdgpu kernel/firmware (Linux).
Long context that causes crashes
The KV cache is added to the weights in memory. Reduce num_ctx or the model size if you approach the ~90 GB limit.

#Machines featuring Strix Halo (August 2026)

The chip's success is reflected in its catalog: more than a dozen machines already use it, from aggressive mini PCs to branded workstations, portable consoles, and modular boxes. The indicative prices below are those observed in August 2026 for 128 GB configurations—they change quickly.

Ryzen AI Max+ 395 machines available or announced (indicative August 2026 prices, 128 GB configuration)
MachineFormatSpecial featureIndicative price
Bosgame M5Mini-PCThe least expensive in the segment~1 700 $
GMKtec EVO-X2Mini-PC4 8K outputshighly variable price, verify it
Beelink GTR9 ProMini-PCDual 10 GbE~2 000 $
Framework DesktopModular caseRepairable, beloved by developershighly variable price, verify it
Corsair AI Workstation 300SFFEstablished brand, solid warranty2 000-3 400 $
Minisforum MS-S1 MAXMini-PCPCIe x16 slot: the only upgradeable one~2 500 $
HP Z2 Mini G1a2.7 L workstationISV-certified, pro channel (PRO 395)2 400-3 500 $
Geekom A9 MegaMini-PC2 TB of original SSD storage~3 200 $
Peladn YO12.4 L mini PCSold as “deploy a 70B”~1 850 $
Acemagic M1A Pro+Mini-PCUp to 24 TB of storagen.c.
Aoostar NEX395Mini-PC140 W TDP (the most powerful)China, EU coming soon
GPD Win 5Portable consoleA 70B in your handsannounced
Minix ER939-AIMini-PCDual 10 GbE variantn.c.

Our buying advice fits in one line: with identical specs (the chip is the same everywhere), use price, warranty, and cooling as the tie-breakers. Tests of each machine are gradually being added to our Reviews & Tests section.

#Frequently asked questions

Can the 128 GB of memory really be used by an LLM?+
For the most part, yes. Memory is unified: under Windows, the GPU can be allocated up to 96 GB, and even more under Linux with the right settings. A 70B model quantized in Q4 (~40 GB) therefore fits entirely in GPU memory, with room for the context.
What is the largest model that can run on Strix Halo?+
A quantized 70B runs comfortably. AMD has demonstrated up to ~120 billion parameters via LM Studio by pushing memory allocation. Beyond that, the limit is no longer capacity but speed: 256 GB/s of bandwidth makes generation slow on very large dense models. MoE models such as GLM or Qwen3-30B-A3B are particularly well suited: few active parameters, many parameters loaded.
Can Strix Halo replace a RTX 5090?+
No—they are not playing the same game. The RTX 5090 is much faster for anything that fits within its 32 GB. Strix Halo wins as soon as the model exceeds that limit: it runs what the 5090 simply cannot load, at a complete-system price lower than the price of the card alone.
Which Strix Halo machine should you choose?+
Since the chip is identical across all of them, the real criteria are price (Bosgame M5 is the most aggressive), warranty and the quality of support (Corsair, HP, and Framework inspire confidence), upgradeability (the Minisforum MS-S1 MAX is the only one with a PCIe x16 slot), and cooling. Read our machine-by-machine reviews before buying.
Windows or Linux for local AI on Strix Halo?+
Both work. Linux offers the most generous GPU memory allocation and the best performance with llama.cpp/ROCm; Windows is simpler with LM Studio, which supports the chip natively. Our setup guide covers both systems, section by section.

#Go further

The Ryzen AI Max+ 395 is part of a broader look at hardware and models for local AI. These site guides build on this one:

Choose your GPU for local AI
The buying guide that puts Strix Halo in perspective against RTX and Mac, based on your budget and target models.
Ollama with AMD GPUs (ROCm)
The detailed ROCm configuration, useful for getting the most out of the Radeon 8060S once you have the platform in hand.
Choose your quantization (Q4, Q5, Q8, FP16)
For choosing between fast Q4 and near-lossless Q8 when memory is plentiful—exactly the case with 128 GB.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.