Ryzen AI Max PRO 400: What 192GB Actually Means for Local LLMs
AMD's new chip advertises 192GB of unified memory and 160GB of VRAM. Here is what that buys you, what it does not, and why memory bandwidth is the number nobody puts on the slide.
Key takeaways
- 192GB total, 160GB for the GPU — the 160GB AMD assigns to the GPU is carved out of the 192GB pool, not added on top. There is no configuration with 192GB for the CPU plus 160GB more for the GPU.
- AMD's own spec sheet states the flagship chip "supports up to 160 GB dedicated graphics memory which is capable of running 300 billion+ parameters at 4-bit quantization."
- That is a capacity claim, not a speed claim. At AMD's stated 273 GB/s bandwidth, a 300B model at 4-bit works out to roughly 1-2 tokens/second by the standard bandwidth-over-weight-size estimate — too slow for interactive chat.
- Three SKUs ship with identical 192GB memory: the Ryzen AI Max+ PRO 495, PRO 490, and PRO 485, differing only in CPU cores and GPU compute units.
- AMD lists availability "in the third quarter of 2026" through OEM partners including HP and Lenovo; this is a commercial/workstation launch, not yet a consumer mini-PC line like the 395 generation.
What AMD actually announced
AMD's Ryzen AI Max PRO 400 series is built on Zen 5 cores with RDNA 3.5 graphics and an XDNA 2 NPU, aimed at commercial PCs, mobile workstations and small-form-factor desktops. On its official blog, AMD states the processors "combine AMD RDNA 3.5 graphics with an AMD XDNA 2 NPU, 192GB of system memory and 160GB of VRAM." That figure comes directly from the manufacturer, not from an independent bench.
The stated target is professional: AI developers, engineers and creators running simulation, content creation and data-heavy workflows. AMD names HP and Lenovo as OEM partners, and hardware press outlet ServeTheHome adds ASUS to the confirmed-partner list.
192GB and 160GB are the same pool, not two pools
The single most confusing part of the announcement: the 160GB assigned to the GPU is not additional to the 192GB total. It is carved out of it, leaving roughly 32GB for the CPU and the OS. Hardware site ServeTheHome describes it as allocating "160GB of their memory pool to the GPU (leaving the last 32GB for the CPU)." There is no way to read the spec as 192GB for the CPU plus a separate 160GB block for the GPU — it is one unified 192GB pool, split.
Do not add the numbers. 192GB + 160GB is not 352GB. Every gigabyte the GPU uses comes out of the same 192GB the system reports as total memory.
What 160GB of GPU memory actually fits
Using the rough Q4 weight sizes this site tracks across model families (roughly 40GB for a 70B dense model, 19-20GB for a 32B model), 160GB leaves substantial headroom above 70B before hitting a ceiling. AMD's own footnote goes further, stating the flagship 495 PRO "supports up to 160 GB dedicated graphics memory which is capable of running 300 billion+ parameters at 4-bit quantization."
| Model size | Estimated Q4 weights | Headroom left in 160GB for context |
|---|---|---|
| 70B dense | ~40GB | ~120GB |
| 120B dense | ~68GB | ~92GB |
| 235B MoE (Qwen3-235B class) | ~130GB | ~30GB |
| 300B (AMD's stated figure) | ~160GB | near zero |
The takeaway from that table: loading a 300B model at 4-bit consumes almost the entire GPU allocation, leaving little room for KV cache. A long context window on a model that size is not guaranteed just because the weights technically fit.
The number the announcement leaves out: bandwidth
Memory capacity is not generation speed. Token throughput on this kind of unified-memory hardware is, in almost every case, limited by memory bandwidth — the model has to be re-read from memory for every token produced. The standard estimate, presented here as a theoretical ceiling rather than a measurement: tokens/second ≈ bandwidth (GB/s) ÷ active weight size (GB).
ServeTheHome reports the PRO 400's LPDDR5X-8533 memory running at 273GB/s. Applying the formula to a 300B model at 4-bit (~160GB of weights): 273 ÷ 160 ≈ 1.7 tokens/second. That is a theoretical ceiling, not a published benchmark from AMD or a reviewer — but it is a low enough number to rule out comfortable interactive chat on the largest model the chip can technically load. A 70B dense model (~40GB) lands closer to 6.8 tokens/second by the same estimate; a 32B model (~20GB) around 13.6 tokens/second.
The gradient nobody puts on the slide. A manufacturer's "up to 192GB" claim is about capacity, never about interactive throughput. On modest-bandwidth hardware like this, MoE models — which activate only a fraction of their parameters per token — get far more usable speed out of a large memory pool than a dense model of the same size.
The three PRO 400 chips
| Model | Cores / threads | Boost clock | GPU compute units | Unified memory |
|---|---|---|---|---|
| Ryzen AI Max+ PRO 495 | 16C / 32T | up to 5.2 GHz | 40 CU (Radeon 8065S) | 192GB |
| Ryzen AI Max PRO 490 | 12C / 24T | up to 5.0 GHz | 32 CU (Radeon 8050S) | 192GB |
| Ryzen AI Max PRO 485 | 8C / 16T | up to 5.0 GHz | 32 CU (Radeon 8050S) | 192GB |
All three ship with the same 192GB memory ceiling; only CPU cores, clocks and GPU compute units change. For local LLM work, the 495 PRO's extra compute units help with prompt processing, but memory bandwidth — and therefore generation speed — is effectively the same across the lineup.
Availability: confirmed versus still unknown
AMD states the series "will be available from leading OEM partners including HP and Lenovo in the third quarter of 2026." ServeTheHome's reporting adds ASUS to the confirmed OEM list. As of this writing, no verified US retail pricing or specific consumer-facing SKU has been confirmed: the initial positioning is commercial and workstation-class, not the budget mini-PC category the 395 generation reached.
Before you plan a purchase: treat any price or ship-date you see outside AMD's own materials or its confirmed OEM partners as unverified. This is a professional-channel launch first.
How it compares to the Ryzen AI Max+ 395
The prior generation, the Ryzen AI Max+ 395 (Strix Halo), tops out at 128GB of unified memory on LPDDR5X-8000, with 256GB/s of theoretical bandwidth on its 256-bit bus — the chip already shipping in consumer mini-PCs. The PRO 400 adds 64GB of memory (192 versus 128) and a modest theoretical bandwidth bump (273 versus 256 GB/s, about 7%), but for now targets commercial systems rather than the budget mini-PC form factor the 395 popularized.
Buy now or wait?
If the goal is running dense 70B-120B models with comfortable context and no multi-GPU workstation, the PRO 400's memory capacity is a real, AMD-confirmed advantage. If the goal is interactive chat on a 300B model, the bandwidth math above (roughly 1-2 tokens/second) argues for waiting on independent, measured benchmarks before committing a professional budget: the capacity is real, the usable speed is not yet demonstrated.
- Already own a 395-based machine (128GB): the extra 64GB rarely justifies an upgrade for typical 30B-70B workloads unless you specifically need very long context or models above 120B.
- Targeting large MoE models: this is the profile that benefits most from a large, bandwidth-limited memory pool, since fewer parameters activate per token.
- Need a consumer machine now: the PRO 400 launches into commercial and workstation channels; confirm a specific configuration fits your budget rather than waiting for a cheaper variant that has not been announced.
For GPU-based alternatives at lower memory but higher bandwidth, see how VRAM sizing works in general, and compare against the existing NVIDIA DGX Spark and the prior-generation Ryzen AI Max+ 395 mini PCs for other unified-memory approaches.
Sources: AMD's official announcement (amd.com) and hardware press analysis from ServeTheHome.
Frequently asked questions
Does the Ryzen AI Max PRO 400 have 192GB or 352GB of AI memory?
192GB total. The 160GB AMD describes as available to the GPU is carved out of that same 192GB pool, not added on top — roughly 32GB remains for the CPU and OS.
Can it really run a 300 billion parameter model?
AMD states 160GB is enough to load a model above 300B parameters at 4-bit quantization. That is a loading-capacity claim from the manufacturer, not a measured speed: at 273GB/s bandwidth, the theoretical throughput estimate is around 1-2 tokens/second, too slow for live conversation.
When will the Ryzen AI Max PRO 400 be available?
AMD lists availability in the third quarter of 2026 through OEM partners including HP and Lenovo, with ASUS confirmed by hardware press. No verified US pricing or specific consumer SKU was confirmed as of this writing.
How does it compare to the Ryzen AI Max+ 395 already on the market?
The 395 tops out at 128GB of unified memory and 256GB/s of theoretical bandwidth. The PRO 400 raises that to 192GB and about 273GB/s, but currently targets commercial and workstation systems rather than the budget mini-PCs the 395 shipped in.
Is memory capacity or memory bandwidth more important for local LLMs?
Both matter for different things: capacity decides which models can be loaded at all, bandwidth decides how fast tokens are generated once a model is loaded. A large model on slow memory still runs — just slowly.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.