Unified Memory Explained: Why It Matters for Running AI Locally
Unified memory puts the CPU and GPU on one shared pool of fast RAM. It is why a Mac or a Strix Halo mini PC can load models no gaming GPU can hold, and also why they generate text more slowly.
Key takeaways
- Unified memory is a single pool of RAM that the CPU and GPU both use directly, instead of separate system RAM and graphics VRAM.
- Its advantage for AI is capacity: up to 128 GB, or more on workstation chips, available to the GPU. The largest consumer graphics card has 32 GB.
- Its weakness is bandwidth: 120 to about 600 GB/s on most chips, against 900 to 1,800 GB/s for high-end GDDR7 cards. Bandwidth sets tokens per second.
- The GPU cannot use all of it. On a Mac the default ceiling is roughly two thirds to three quarters of total memory.
- It is soldered. Whatever you buy on day one is what the machine will have forever, so this is the spec to get right.
What unified memory is
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- 30-day refund
In a conventional PC with a graphics card, there are two memory systems. The CPU works from system RAM on the motherboard. The GPU works from its own VRAM on the card. Anything the GPU needs must be copied across the PCIe bus, which is an order of magnitude slower than either memory. A model that does not fit in VRAM has to be split, with part of it served from system RAM at a painful speed penalty.
A unified memory architecture removes the split. CPU, GPU and any neural accelerator sit on one chip package and address the same physical memory, wired close to the processor with a wide bus. There is no copy: the GPU reads model weights from the same place the CPU loaded them. Apple brought the design to mainstream computers with the M1 in 2020 and describes it in its Mac specifications; AMD's Ryzen AI Max ("Strix Halo") and NVIDIA's DGX Spark apply the same idea on the PC side.
Integrated graphics in ordinary laptops have shared system RAM for decades, so what is new? Width and speed. Ordinary shared memory is a dual-channel bus at 50–100 GB/s. Unified memory designs use four to sixteen channels of LPDDR5X, reaching several hundred GB/s.
Unified memory vs VRAM vs system RAM
| System RAM (DDR5) | Unified memory | VRAM (GDDR6X / GDDR7) | |
|---|---|---|---|
| Used by | CPU | CPU and GPU together | GPU only |
| Typical bandwidth | 60–100 GB/s | 120–600+ GB/s | 500–1,800 GB/s |
| Capacity available to the GPU | Only by slow spill-over | Up to roughly 75% of 16–128 GB and beyond | 8–32 GB on consumer cards |
| Copy needed between CPU and GPU | Yes, over PCIe | No | Yes, over PCIe |
| Upgradeable | Yes | No | No |
| Power draw under AI load | n/a | Tens of watts for the whole system | 200–575 W for the card alone |
The short version: VRAM is the fast, small option; unified memory is the large, moderately fast one; system RAM is the large, slow one. For background on the first, see what is VRAM.
Bandwidth by chip
| Chip | Max unified memory | Memory bandwidth |
|---|---|---|
| Apple M4 | 32 GB | 120 GB/s |
| Apple M4 Pro | 64 GB | 273 GB/s |
| Apple M4 Max | 128 GB | 410 or 546 GB/s depending on the GPU variant |
| Apple M3 Ultra | 512 GB | 819 GB/s |
| Apple M5 Max (40-core GPU) | 128 GB | 614 GB/s |
| AMD Ryzen AI Max+ 395 (Strix Halo) | 128 GB | 256 GB/s |
| NVIDIA DGX Spark (GB10) | 128 GB | 273 GB/s |
| For comparison: RTX 4090 | 24 GB VRAM | 1,008 GB/s |
| For comparison: RTX 5090 | 32 GB VRAM | 1,792 GB/s |
Manufacturer specifications. Note that binned chips with fewer GPU cores often ship with lower bandwidth; check the exact variant.
How much can the GPU actually use?
Not all of it. The operating system and your applications need memory too, and macOS caps the GPU's share by default at roughly two thirds on smaller machines and three quarters on larger ones. The BestLLMfor hardware database uses these working figures:
| Mac configuration | Usable for a model (our database) | Example of what fits at 4-bit |
|---|---|---|
| 16 GB | ≈ 10 GB | Qwen 3 14B (9 GB), tight |
| 24 GB | ≈ 16 GB | gpt-oss 20B (13 GB) |
| 36 GB | ≈ 27 GB | Qwen 3.6 35B-A3B (21 GB) |
| 48 GB | ≈ 36 GB | Qwen 3.6 35B-A3B with a long context |
| 64 GB | ≈ 48 GB | Llama 3.3 70B (40 GB) |
| 128 GB | ≈ 96 GB | gpt-oss 120B (70 GB) |
Model sizes from the BestLLMfor catalog, September 20, 2026.
On macOS the ceiling can be raised from the terminal, at the cost of memory for everything else:
# Allow the GPU to wire up to ~56 GB on a 64 GB Mac (resets at reboot)
sudo sysctl iogpu.wired_limit_mb=57344
Leave at least 8 GB for the system, more if you keep a browser open. On Strix Halo machines the split between system and graphics memory is set in the BIOS or driver, and up to 96 GB can typically be assigned to the GPU on a 128 GB system.
What it means for speed
Generating a token requires reading the active weights once, so the ceiling on tokens per second is bandwidth divided by model size. For a dense 40 GB model such as Llama 3.3 70B at 4-bit:
| Machine | Fits? | Bandwidth | Theoretical ceiling |
|---|---|---|---|
| RTX 5090 (32 GB) | No: spills to system RAM | — | A few tok/s in practice |
| 2 × RTX 3090 (48 GB) | Yes | 936 GB/s per card | ≈ 23 tok/s |
| Mac with M4 Max 64 GB | Yes | 546 GB/s | ≈ 14 tok/s |
| Strix Halo 128 GB | Yes | 256 GB/s | ≈ 6 tok/s |
| M3 Ultra 512 GB | Yes | 819 GB/s | ≈ 20 tok/s |
Upper bounds; real results land below.
Two lessons. First, fitting beats everything: the fastest card on the market loses badly the moment the model overflows. Second, unified memory machines are at their best with mixture-of-experts models, which need lots of capacity but read only a few billion parameters per token. On a 128 GB M5 Max we measured gpt-oss 120B at 75–79 tokens per second, against 37 for a dense 27B model (benchmark). The mechanism is explained in mixture of experts explained.
The weaker side is prompt processing. Reading a long prompt is compute-bound, and integrated GPUs have far fewer compute units than a high-end discrete card. Expect long documents to take noticeably longer to ingest than on an RTX 4090 or 5090, even when generation speed is comparable.
Software support
- Apple Silicon: excellent. llama.cpp (Metal), Ollama, LM Studio and Apple's MLX all use unified memory natively. MLX builds are often the fastest; see MLX on Apple Silicon.
- Strix Halo: good and improving, through Vulkan and ROCm in llama.cpp-based tools. See the Ryzen AI Max+ 395 mini PC page.
- DGX Spark: the full CUDA stack, which is its main selling point; see NVIDIA DGX Spark.
- Not supported on any of them the way it is on discrete NVIDIA cards: much training and research code, which assumes CUDA. Background in what is CUDA.
Who should buy unified memory for AI?
| Your priority | Better choice |
|---|---|
| Models up to about 30B, maximum speed, image generation | A discrete GPU with 16–32 GB of VRAM |
| 70B-class dense models or 100B+ MoE models on one quiet machine | Unified memory, 64–128 GB |
| Portability and battery life | Unified memory (MacBook Pro) |
| Fine-tuning, research code, vLLM | Discrete NVIDIA GPUs |
| Lowest power draw and noise | Unified memory |
If you go the unified route, buy memory before anything else: it cannot be added later, and it is the only spec that decides which models you can run. Compare options in the AI hardware hub and size models with the VRAM calculator. Hardware figures are open through the BestLLMfor public API (CC BY 4.0) and our MCP server. Apple's MLX framework is developed in the open at github.com/ml-explore/mlx, and the llama.cpp project tracks backend support on GitHub.
Frequently asked questions
Is unified memory the same as RAM?
It is RAM, but organized differently: one pool on the chip package, shared directly by the CPU and GPU over a wide, fast bus. A conventional PC has separate system RAM and graphics VRAM connected by the much slower PCIe bus.
Is unified memory better than VRAM for AI?
It offers far more capacity, so larger models fit. High-end VRAM offers two to four times the bandwidth, so models that fit on a graphics card run faster there. Capacity favors unified memory; speed favors VRAM.
How much unified memory do I need for local LLMs?
16 GB runs 8B models; 24 to 36 GB runs 14B to 30B models; 64 GB runs 70B models at 4-bit; 128 GB runs 100B-class mixture-of-experts models. Roughly a quarter to a third of the total is reserved for the system.
Can I upgrade unified memory later?
No. It is soldered to the processor package. Choose the capacity at purchase time, since it determines which models the machine can ever run.
Why is my Mac slower than an RTX card on the same model?
Lower memory bandwidth limits generation speed, and the integrated GPU has less compute for prompt processing. The Mac's advantage appears with models too large for any single graphics card.
Do Windows PCs have unified memory?
Some do. AMD's Ryzen AI Max (Strix Halo) systems and NVIDIA's DGX Spark use a unified design with up to 128 GB. Ordinary laptops with integrated graphics share system RAM too, but over a much narrower, slower bus.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.