VRAM: what it is and how much you need for AI ?
VRAM (Video RAM) is the memory soldered onto the graphics card and reserved for the GPU: 8 to 32 GB on consumer cards in late September 2026, versus up to 512 GB of unified memory on a Mac Studio M5 Ultra. It directly determines which local AI models can load: an LLM that exceeds its capacity spills into RAM and slows down by 5 to 20 times. Capacity determines what fits; bandwidth determines speed.
VRAM, or video memory, is the memory soldered onto your graphics card and reserved for the GPU. It is much faster than the computer's RAM, and it determines which AI models you can run locally: an LLM must fit entirely in it to respond quickly. As of September 20, 2026, consumer cards range from 8 to 32 GB. This page explains what VRAM is, how to find yours in thirty seconds, and what each tier actually enables.
#VRAM in one minute
VRAM stands for Video Random Access Memory. It is memory, like RAM, but located a few millimeters from the graphics chip and connected to it by a very wide bus. Its original role is to store what the GPU constantly manipulates: the textures, geometry, and images of a game. For AI, it stores the model weights and its working memory.
Two numbers define it. Capacity, in gigabytes, tells you what can fit. Bandwidth, in gigabytes per second, tells you how quickly the GPU can read it. In gaming, capacity is the main concern. In AI, both matter: capacity determines whether the model loads, while bandwidth determines how quickly it writes.
#RAM and VRAM: the differences
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
| RAM (system memory) | VRAM (video memory) | |
|---|---|---|
| Location | Memory sticks on the motherboard | Chips soldered onto the graphics card |
| Serving | The processor (CPU) | The graphics processor (GPU) |
| Technology | DDR4, DDR5 | GDDR6, GDDR6X, GDDR7 |
| Bandwidth | 60 to 100 GB/s in dual-channel DDR5 | 360 GB/s (RTX 3060) to 1,792 GB/s (RTX 5090) |
| Typical capacity | 16 to 64 GB | 8 to 32 GB |
| Scalable | Yes, we add memory sticks | No, it's soldered |
The bandwidth gap, on the order of ten to twenty times, explains everything else on this page. An AI model rereads all its weights for every token it produces. From VRAM, this is fast. From RAM, the same model runs five to twenty times slower.
#Know your VRAM
- Windows, without installing anything
- Open Task Manager (Ctrl + Shift + Esc), select the Performance tab, then GPU in the left column. The “Dedicated GPU memory” line shows your VRAM. The “Shared GPU memory” displayed beside it is not VRAM: it is RAM that Windows can lend to the GPU, and it is much slower.
- Windows, alternative method
- Windows key + R, type dxdiag, Display tab: the “Display Memory (VRAM)” line.
- Windows and Linux with a NVIDIA card
- In a terminal, the nvidia-smi command displays total and used memory in real time. It's the tool to keep open while a model is running.
- Linux with an AMD card
- The rocm-smi --showmeminfo vram command, or your distribution’s graphical tool.
- Mac
- Apple menu, About This Mac: the Memory line. A Mac with Apple Silicon has no separate VRAM; see the dedicated section below.
#Why local AI depends on VRAM
An LLM uses VRAM in three ways. First, the weights: the number of parameters multiplied by the precision. With 4-bit quantization, count about 0.6 GB per billion parameters, or 5 GB for an 8B model. Next, the KV cache, which stores the current conversation and grows with the context length. Finally, an operating margin of around 1 to 2 GB.
If the total exceeds VRAM, tools such as Ollama or llama.cpp offload part of the model to RAM: this is CPU+GPU offloading, a documented llama.cpp feature for running models larger than the available VRAM. The model still works, but speed collapses as soon as a few layers cross over, because each token then has to wait for a read from system RAM, which is ten to twenty times slower. That’s why an older card with lots of VRAM, such as a RTX 3090 with 24 GB, remains more useful for local AI than a recent 12 GB card, even though the latter is faster for gaming and 3D rendering.
- VRAM calculator: weights, context, and headroom for your model
- Choose your quantization: Q4, Q5, Q8, or FP16
#How much VRAM for each model
The table cross-references our database of 90 cards and chips with the 249 models in the catalog. For each tier, it lists a representative model that fits entirely in VRAM in Q4_K_M. Keep 1 to 3 GB of headroom for the context: when the model's size approaches the card's capacity, the conversation must remain short.
| VRAM | Typical cards | Representative model that fits | Weights in Q4 | Detailed guide |
|---|---|---|---|---|
| 6 GB | RTX 2060, GTX 1660, RTX 4050 portable | Qwen 3 8B, for short context | 5 GB | Which LLM for 6 GB |
| 8 GB | RTX 3060 Ti, RTX 4060, RTX 5060 | Qwen 3 8B at ease, Gemma 4 12B with a short context | 5 and 7 GB | Which LLM for 8 GB |
| 12 GB | RTX 3060 12 GB, RTX 4070, RTX 5070 | Qwen 3 14B | 9 GB | Which LLM for 12 GB |
| 16 GB | RTX 4060 Ti 16 GB, RTX 5070 Ti, RTX 5080, RX 9070 XT | gpt-oss 20B, Mistral Small 3.2 24B | 13 and 14 GB | Which LLM for 16 GB |
| 24 GB | RTX 3090, RTX 4090, RX 7900 XTX | Qwen 3.8 27B, Qwen 3.6 35B-A3B | 16 and 21 GB | Which LLM for 24 GB |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B in Q5 | 25 GB | Which LLM for 32 GB |
| 48 GB | 2 × RTX 3090 or 4090, 64 GB Mac | Llama 3.3 70B | 40 GB | View the calculator |
| 96 GB | Mac 128 GB | gpt-oss 120B | 70 GB | View the calculator |
- Which LLM for 8 GB of VRAM?
- Which LLM for 12 GB of VRAM?
- Which LLM for 16 GB of VRAM?
- Which LLM for 24 GB of VRAM?
#Can you increase your VRAM?
On a dedicated graphics card, no, regardless of the manufacturer. The memory chips are soldered and sized with the GPU during design: the only way to get more VRAM is to replace the card or add a second one, which local AI tools can use by automatically distributing the model across the cards. Settings that promise to “increase VRAM” in Windows or the BIOS do not create additional memory: they concern two special cases, described below, that provide no benefit to a dedicated card already fully recognized by the system.
- Integrated GPU (iGPU)
- It has no dedicated VRAM and borrows some system RAM. The BIOS often lets you increase this allocation (the UMA Frame Buffer Size setting or equivalent). This increases capacity, not speed, which remains that of RAM.
- Windows shared GPU memory
- Windows allows a dedicated card to spill over into RAM. For a game, this prevents a crash. For an LLM, it is the slow scenario described above.
The real room for maneuver is software: choose a lighter quantization, reduce the context window, or quantize the KV cache. These three levers often save several gigabytes without changing the hardware.
#The Mac case: unified memory
Macs with Apple Silicon do not have separate VRAM. The CPU and GPU share a single, unified memory pool with high bandwidth. In the M4 generation: 120 GB/s for the base chip, up to 546 GB/s for an M4 Max. The GPU cannot claim all of that memory because macOS reserves part of it; in practice, plan for about 16 GB usable by a model on a 24 GB Mac, 48 GB on a 64 GB Mac, and 96 GB on a 128 GB Mac—a reserve that weighs proportionally more heavily on smaller configurations.
The lineup has changed since then: Apple launched the M5 chip in late 2025 (153.6 GB/s, 32 GB maximum memory), followed by the M5 Pro (307 GB/s, 64 GB) and the M5 Max (up to 614 GB/s, 128 GB). The Mac Studio followed with the M5 Ultra, announced on August 25, 2026: Apple advertises “up to 512 GB” of unified memory and “1.2 TB/s” of bandwidth, theoretically enough to load a model with several hundred billion parameters. That same day, Apple released the M6, built on a 2 nm process and available since September 22, 2026: it is only an entry-level chip, with no Pro or Max variant this generation, capped at 32 GB of memory and 170 GB/s of bandwidth—less than an M5 Pro released a year earlier. One point to check before buying a “latest-generation” Mac for AI: the chipset name does not tell the whole story; the variant matters more than the model year.
| Chip | Bandwidth | Maximum memory |
|---|---|---|
| M4 | 120 GB/s | 32 GB |
| M4 Pro | 273 GB/s | 64 GB |
| M4 Max | 410 to 546 GB/s | 128 GB |
| M5 | 153.6 GB/s | 32 GB |
| M5 Pro | 307 GB/s | 64 GB |
| M5 Max | up to 614 GB/s | 128 GB |
| M5 Ultra | 1.2 TB/s (1,228.8 GB/s) | 512 GB |
| M6 | 170 GB/s | 32 GB |
This is what makes high-end Macs distinctive for local AI: no consumer graphics card offers 96 GB of usable VRAM, let alone the hundreds of gigabytes in a Mac Studio M5 Ultra. The trade-off is that, at equal capacity, a recent NVIDIA card is still faster thanks to higher bandwidth per gigabyte: the 1,792 GB/s of a RTX 5090 with 32 GB far exceeds the 614 GB/s of an M5 Max with 128 GB, but the latter can load a model four times larger.
- Hardware for local AI: our machine guides
- Which GPU for a local LLM?
- Source: Apple Newsroom, M6 and M5 Ultra (25/08/2026)
- Source: NVIDIA GeForce RTX 5090 specification sheet
- Source: llama.cpp GitHub repository (multi-GPU distribution)
#FAQ
What’s the difference between RAM and VRAM?+
Is 8 GB of VRAM enough in 2026?+
Why does Windows show more GPU memory than my card has?+
Do two graphics cards combine their VRAM?+
Does VRAM matter more than GPU power for local AI?+
Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — 16 GB of VRAM, enough to load a 14B model quantized to Q4 with its context. All AI hardware →
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.