Beginner 11 minConcepts

VRAM: what it is and how much you need for AI ?

Direct response

VRAM (Video RAM) is the memory soldered onto the graphics card and reserved for the GPU: 8 to 32 GB on consumer cards in late September 2026, versus up to 512 GB of unified memory on a Mac Studio M5 Ultra. It directly determines which local AI models can load: an LLM that exceeds its capacity spills into RAM and slows down by 5 to 20 times. Capacity determines what fits; bandwidth determines speed.

VRAM, or video memory, is the memory soldered onto your graphics card and reserved for the GPU. It is much faster than the computer's RAM, and it determines which AI models you can run locally: an LLM must fit entirely in it to respond quickly. As of September 20, 2026, consumer cards range from 8 to 32 GB. This page explains what VRAM is, how to find yours in thirty seconds, and what each tier actually enables.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#VRAM in one minute

VRAM stands for Video Random Access Memory. It is memory, like RAM, but located a few millimeters from the graphics chip and connected to it by a very wide bus. Its original role is to store what the GPU constantly manipulates: the textures, geometry, and images of a game. For AI, it stores the model weights and its working memory.

Two numbers define it. Capacity, in gigabytes, tells you what can fit. Bandwidth, in gigabytes per second, tells you how quickly the GPU can read it. In gaming, capacity is the main concern. In AI, both matter: capacity determines whether the model loads, while bandwidth determines how quickly it writes.

#RAM and VRAM: the differences

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Bandwidths: manufacturer specifications
RAM (system memory)VRAM (video memory)
LocationMemory sticks on the motherboardChips soldered onto the graphics card
ServingThe processor (CPU)The graphics processor (GPU)
TechnologyDDR4, DDR5GDDR6, GDDR6X, GDDR7
Bandwidth60 to 100 GB/s in dual-channel DDR5360 GB/s (RTX 3060) to 1,792 GB/s (RTX 5090)
Typical capacity16 to 64 GB8 to 32 GB
ScalableYes, we add memory sticksNo, it's soldered

The bandwidth gap, on the order of ten to twenty times, explains everything else on this page. An AI model rereads all its weights for every token it produces. From VRAM, this is fast. From RAM, the same model runs five to twenty times slower.

#Know your VRAM

Windows, without installing anything
Open Task Manager (Ctrl + Shift + Esc), select the Performance tab, then GPU in the left column. The “Dedicated GPU memory” line shows your VRAM. The “Shared GPU memory” displayed beside it is not VRAM: it is RAM that Windows can lend to the GPU, and it is much slower.
Windows, alternative method
Windows key + R, type dxdiag, Display tab: the “Display Memory (VRAM)” line.
Windows and Linux with a NVIDIA card
In a terminal, the nvidia-smi command displays total and used memory in real time. It's the tool to keep open while a model is running.
Linux with an AMD card
The rocm-smi --showmeminfo vram command, or your distribution’s graphical tool.
Mac
Apple menu, About This Mac: the Memory line. A Mac with Apple Silicon has no separate VRAM; see the dedicated section below.
Monitor VRAM while a model is running (NVIDIA)
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv -l 2

#Why local AI depends on VRAM

An LLM uses VRAM in three ways. First, the weights: the number of parameters multiplied by the precision. With 4-bit quantization, count about 0.6 GB per billion parameters, or 5 GB for an 8B model. Next, the KV cache, which stores the current conversation and grows with the context length. Finally, an operating margin of around 1 to 2 GB.

If the total exceeds VRAM, tools such as Ollama or llama.cpp offload part of the model to RAM: this is CPU+GPU offloading, a documented llama.cpp feature for running models larger than the available VRAM. The model still works, but speed collapses as soon as a few layers cross over, because each token then has to wait for a read from system RAM, which is ten to twenty times slower. That’s why an older card with lots of VRAM, such as a RTX 3090 with 24 GB, remains more useful for local AI than a recent 12 GB card, even though the latter is faster for gaming and 3D rendering.

#How much VRAM for each model

The table cross-references our database of 90 cards and chips with the 249 models in the catalog. For each tier, it lists a representative model that fits entirely in VRAM in Q4_K_M. Keep 1 to 3 GB of headroom for the context: when the model's size approaches the card's capacity, the conversation must remain short.

Calculated from the QuelLLM catalog and hardware database · Q4_K_M weights · 20/09/2026
VRAMTypical cardsRepresentative model that fitsWeights in Q4Detailed guide
6 GBRTX 2060, GTX 1660, RTX 4050 portableQwen 3 8B, for short context5 GBWhich LLM for 6 GB
8 GBRTX 3060 Ti, RTX 4060, RTX 5060Qwen 3 8B at ease, Gemma 4 12B with a short context5 and 7 GBWhich LLM for 8 GB
12 GBRTX 3060 12 GB, RTX 4070, RTX 5070Qwen 3 14B9 GBWhich LLM for 12 GB
16 GBRTX 4060 Ti 16 GB, RTX 5070 Ti, RTX 5080, RX 9070 XTgpt-oss 20B, Mistral Small 3.2 24B13 and 14 GBWhich LLM for 16 GB
24 GBRTX 3090, RTX 4090, RX 7900 XTXQwen 3.8 27B, Qwen 3.6 35B-A3B16 and 21 GBWhich LLM for 24 GB
32 GBRTX 5090Qwen 3.6 35B-A3B in Q525 GBWhich LLM for 32 GB
48 GB2 × RTX 3090 or 4090, 64 GB MacLlama 3.3 70B40 GBView the calculator
96 GBMac 128 GBgpt-oss 120B70 GBView the calculator
→
The pre-purchase check
At the same budget, for local AI, prioritize VRAM capacity over the GPU generation. Going from 12 to 16 GB opens up the class of 20- to 24-billion-parameter models; going from 16 to 24 GB opens up the 27- to 35-billion class. No speed gain makes up for a model that will not load.

#Can you increase your VRAM?

On a dedicated graphics card, no, regardless of the manufacturer. The memory chips are soldered and sized with the GPU during design: the only way to get more VRAM is to replace the card or add a second one, which local AI tools can use by automatically distributing the model across the cards. Settings that promise to “increase VRAM” in Windows or the BIOS do not create additional memory: they concern two special cases, described below, that provide no benefit to a dedicated card already fully recognized by the system.

Integrated GPU (iGPU)
It has no dedicated VRAM and borrows some system RAM. The BIOS often lets you increase this allocation (the UMA Frame Buffer Size setting or equivalent). This increases capacity, not speed, which remains that of RAM.
Windows shared GPU memory
Windows allows a dedicated card to spill over into RAM. For a game, this prevents a crash. For an LLM, it is the slow scenario described above.

The real room for maneuver is software: choose a lighter quantization, reduce the context window, or quantize the KV cache. These three levers often save several gigabytes without changing the hardware.

#The Mac case: unified memory

Macs with Apple Silicon do not have separate VRAM. The CPU and GPU share a single, unified memory pool with high bandwidth. In the M4 generation: 120 GB/s for the base chip, up to 546 GB/s for an M4 Max. The GPU cannot claim all of that memory because macOS reserves part of it; in practice, plan for about 16 GB usable by a model on a 24 GB Mac, 48 GB on a 64 GB Mac, and 96 GB on a 128 GB Mac—a reserve that weighs proportionally more heavily on smaller configurations.

The lineup has changed since then: Apple launched the M5 chip in late 2025 (153.6 GB/s, 32 GB maximum memory), followed by the M5 Pro (307 GB/s, 64 GB) and the M5 Max (up to 614 GB/s, 128 GB). The Mac Studio followed with the M5 Ultra, announced on August 25, 2026: Apple advertises “up to 512 GB” of unified memory and “1.2 TB/s” of bandwidth, theoretically enough to load a model with several hundred billion parameters. That same day, Apple released the M6, built on a 2 nm process and available since September 22, 2026: it is only an entry-level chip, with no Pro or Max variant this generation, capped at 32 GB of memory and 170 GB/s of bandwidth—less than an M5 Pro released a year earlier. One point to check before buying a “latest-generation” Mac for AI: the chipset name does not tell the whole story; the variant matters more than the model year.

Apple Silicon chips: unified memory and bandwidth (M4 to M6 comparison)
ChipBandwidthMaximum memory
M4120 GB/s32 GB
M4 Pro273 GB/s64 GB
M4 Max410 to 546 GB/s128 GB
M5153.6 GB/s32 GB
M5 Pro307 GB/s64 GB
M5 Maxup to 614 GB/s128 GB
M5 Ultra1.2 TB/s (1,228.8 GB/s)512 GB
M6170 GB/s32 GB

This is what makes high-end Macs distinctive for local AI: no consumer graphics card offers 96 GB of usable VRAM, let alone the hundreds of gigabytes in a Mac Studio M5 Ultra. The trade-off is that, at equal capacity, a recent NVIDIA card is still faster thanks to higher bandwidth per gigabyte: the 1,792 GB/s of a RTX 5090 with 32 GB far exceeds the 614 GB/s of an M5 Max with 128 GB, but the latter can load a model four times larger.

#FAQ

What’s the difference between RAM and VRAM?+
RAM serves the processor and is located on the motherboard; VRAM serves the GPU and is soldered to the graphics card. VRAM has ten to twenty times more memory bandwidth, but less capacity and no upgrade path. An AI model that fits in VRAM responds much faster than one that spills over into RAM.
Is 8 GB of VRAM enough in 2026?+
For 1080p gaming, yes in most cases. For local AI, 8 GB supports models up to 12 billion parameters in Q4, such as Gemma 4 12B, provided you keep the context short; an 8B model is more comfortable. It is a good starting point, but quickly becomes limiting if you are targeting models of 20 billion parameters or more.
Why does Windows show more GPU memory than my card has?+
Task Manager adds together two different figures: dedicated memory, your actual VRAM soldered onto the card, and shared memory, a portion of system RAM that Windows allows the GPU to borrow when needed. Only dedicated memory counts when assessing what an AI model can load at full speed: as soon as an LLM spills into shared memory, generation slows dramatically, exactly as if it were running on the CPU.
Do two graphics cards combine their VRAM?+
For AI, yes: inference engines such as llama.cpp, Ollama, or vLLM can distribute a model’s layers across multiple GPUs, so two 24 GB cards can load a model weighing up to 40–45 GB in practice. For gaming, however, the answer is no: each card continues to use its own memory, with no way to combine them.
Does VRAM matter more than GPU power for local AI?+
To run a model, yes: capacity determines what can be loaded into memory, and memory bandwidth then determines token-by-token generation speed, much more than the GPU’s raw computing power. The latter matters primarily for reading the initial prompt and generating images, two tasks that are much more compute-intensive than repeated memory reads.

Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — 16 GB of VRAM, enough to load a 14B model quantized to Q4 with its context. All AI hardware →

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.