Which LLM for 32 GB of VRAM ?
With 32 GB of VRAM, you can fully load a dense 27B to 32B model in Q4 (16 to 20 GB) with a long context, a 30-35B MoE (19 to 21 GB), or gpt-oss-20b (14 GB) with plenty of headroom. The threshold that does not fit is 70B: 40 to 43 GB in Q4, so it necessarily spills partly into system RAM, where speed collapses. Always account for the context cache in addition to the weights.
32 GB of VRAM is the threshold where you stop choosing between model size and context length. This page explains what really fits, what changes depending on the card (RTX 5090 or Radeon AI PRO R9700), why a 70B or 123B does not become comfortable, and when a Mac’s unified memory or a second card is a better purchase.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#32 GB of VRAM: what the capacity changes compared with 24 GB
A Q4 model uses approximately 0.6 GB per billion parameters, plus the context cache, which grows with conversation length. At 24 GB, a 27–32B dense model in Q4 (16–20 GB) fits, but the context must remain short. At 32 GB, 12–16 GB remains for context, activations, and the system: this is the budget that makes 32,000- to 128,000-token windows possible on a 27B, or a 30–35B MoE with comfortable headroom.
| Model | Q4 weight | Remaining capacity for context and system | Verdict |
|---|---|---|---|
| gpt-oss-20b | 14 GB | 18 GB | Very wide margin, long context |
| Qwen 3.5 27B / Qwen 3.8 27B | 16 GB | 16 GB | Comfortable, long context possible |
| Gemma 4 31B | 18 GB | 14 GB | Comfortable |
| Qwen 3 32B | 19–20 GB | 12 GB | Good, context to monitor |
| Qwen3-Coder 30B-A3B (MoE) | 19 GB | 13 GB | Good, very fast for its size |
| Qwen 3.6 35B-A3B (MoE) | 21 GB | 11 GB | Good, medium context |
| Llama 3.3 70B | 43 GB | negative | Doesn't fit: offloading required |
These values are file weights, not total requirements: the site’s calculator adds the cache for a given model and context. The cache can also be quantized to q8_0, which, according to Ollama’s FAQ, cuts its size by about half compared with f16, provided Flash Attention is enabled.
#Which cards offer 32 GB, and why bandwidth matters
Two desktop cards stand out. The GeForce RTX 5090 includes 32 GB of GDDR7 on a 512-bit bus, with 1,792 GB/s of bandwidth according to NVIDIA. AMD's Radeon AI PRO R9700 also offers 32 GB, but in GDDR6 on a 256-bit bus at 640 GB/s, according to AMD, for desktop and professional use. Capacity is identical; generation speed is not.
| Card | Memory | Bandwidth | Ceiling for a 19 GB model |
|---|---|---|---|
| GeForce RTX 5090 | 32 GB GDDR7, 512-bit | 1,792 GB/s | approximately 94 t/s |
| Radeon AI PRO R9700 | 32 GB GDDR6, 256-bit | 640 GB/s | approximately 34 t/s |
The ceiling is calculated by dividing bandwidth by the weights read for each token; it is never fully achieved, but the ratio remains valid: with the same model, the RTX 5090 generates more than twice as fast as the R9700. For an AMD card, the driver and software stack (ROCm, Vulkan) also matter; the dedicated guide to AMD cards details the limitations. Specific cards have their own pages: the RTX 5090 page provides the card's details.
Two other approaches reach 32 GB. Two 16 GB cards can split a model with llama.cpp, but the PCIe connection slows communication and the model must be split cleanly; the multi-GPU guide provides the method. A Mac with 32 GB or more of unified memory is a third option, with lower bandwidth but no dedicated VRAM limit.
Prices change quickly: every Monday and Thursday, the site’s tracker records the lowest price for graphics cards for local AI, along with the price per GB of VRAM.
#Models to prioritize on 32 GB
- Dense 27B to 32B
- The quality choice: Qwen 3.5 27B or 3.8 27B (16 GB in Q4), Gemma 4 31B (18 GB), Qwen 3 32B (19 to 20 GB). They leave room for a long context and work well with finer quantizations (Q5, Q6) if you want higher quality.
- MoE 30–35B
- Qwen3-Coder 30B-A3B (19 GB) and Qwen 3.6 35B-A3B (21 GB) activate only about 3 billion parameters per token: they offer the best speed for their size and are the right choice for code and agents.
- gpt-oss-20b
- 14 GB in the Ollama library; the Ollama page says it can run on systems with at least 16 GB of memory. On 32 GB, it leaves room for a huge context and lets you keep a second model loaded at the same time.
- 24B code model
- Devstral Small 2 24B (14 GB): a specialist in coding agents that leaves room for two resident models.
To choose, ask three questions. Does the model need to handle a long context (documents, code repository)? Choose a 27B dense model in Q4, leaving 16 GB for cache. Do you need speed (agents, tool loops)? Choose the MoE. Do you simply want more quality than at 24 GB? Move from Q4 to Q6 on a 27B to 32B model: there's room.
#Verify that the model really fits on the card
A model that “loads” is not necessarily entirely on the GPU. Ollama silently distributes layers between the card and RAM when VRAM is insufficient, and speed drops without an error message. The official FAQ states that the ollama ps command displays where the model was loaded in the Processor column: 100% GPU means it is entirely on the GPU, while a split such as 30%/70% CPU/GPU indicates overflow.
- 01Load the model with the target contextRun the model with the actual context size of your use case, not the default value: the context cache is what causes overflow.
- 02Read the Processor columnRun ollama ps in another terminal. Aim for 100% GPU; any CPU share cuts the speed.
- 03Reduce if necessaryIf the model runs out of room, reduce the context, enable q8_0 cache quantization with Flash Attention, or switch to a lighter quantization.
- 04Check remaining VRAMOn Windows or Linux, nvidia-smi shows used memory; leave one to two GB of headroom for the display and spikes.
- 05Measure on your workloadCompare throughput with a short prompt and with your real prompt: the second number is what matters.
#70B and 123B: why offloading does not make these models usable
Llama 3.3 70B weighs 43 GB in the Ollama library, Mistral Large 123B 73 GB. On 32 GB, you therefore need to leave at least 11 GB for the 70B and 41 GB for the 123B in system memory. For each token, the processor reads these layers at system memory speed: about 102 GB/s at most for dual-channel DDR5-6400 (6,400 MT/s × 8 bytes × 2). Dividing by 41 GB gives the 123B a ceiling of about 2.5 tokens per second; for the 70B, with 11 GB in RAM, about 9 tokens per second. These are ceilings, not measurements: they mainly show that throughput does not degrade progressively, but in steps, as soon as the layers leave the card.
| Model | Size Ollama | Moves to system RAM | RAM ceiling (102 GB/s) |
|---|---|---|---|
| Llama 3.3 70B | 43 GB | about 11 GB (34%) | about 9 t/s |
| Mistral Large 123B | 73 GB | approximately 41 GB (56%) | about 2.5 t/s |
This calculation assumes the card keeps 32 GB of weights; in practice, the context cache also takes up space, so more spills into RAM and the limit drops. MoEs behave better: llama.cpp offers an option to keep experts from certain layers on the CPU (--n-cpu-moe), for configurations that cannot load the entire model on the GPU. It saves a model that doesn't fit, but a llama.cpp repository user reports an approximately 80% speed loss when putting all experts on the CPU: it's a last-resort solution.
The right value for --n-cpu-moe depends on your memory; you have to find it through testing. For a dense 70B, the right answer at 32 GB is almost always a smaller model or more memory, not offload.
#Fine-tuning a model on 32 GB: what works and what doesn't
QLoRA fine-tuning keeps the model weights in 4-bit form and trains only small adapters. The weights alone for a 70B model already take up 40 GB, so QLoRA on a 70B model does not fit in 32 GB. A 32B model in 4-bit form takes up about 19 to 20 GB: that leaves 12 GB for activations, the optimizer, and the adapters, which is feasible with a short context and a batch size of 1, with no safety margin. A 7B to 14B model is comfortable. These are rough estimates to confirm in the tool you use.
#32 GB of VRAM or another solution: the decision table
| Need | Solution | Why |
|---|---|---|
| Dense 27–32B, long context, maximum speed | 32 GB card NVIDIA | High bandwidth, CUDA ecosystem |
| Same capacity at lower electrical power | Radeon AI PRO R9700 | Same 32 GB, lower theoretical bandwidth |
| 64 GB and larger models, occasional use | Mac with 64 GB or more of unified memory | Capacity without offload, lower speed |
| Limited budget, 30B MoE | A 24 GB card is often enough | The 19 GB MoE fits with a moderate context |
- Which LLM for 24 GB of VRAM
- Mac Studio for 64 to 512 GB
- Source: NVIDIA, announcement of the RTX 50
- Source: AMD, Radeon AI PRO R9700
- Source: Ollama library, Mistral Large
#Frequently asked questions
What is the best LLM for 32 GB of VRAM?+
Is 32 GB of VRAM enough for Mistral Large 123B?+
RTX 5090 or Mac Studio 64 GB for local AI?+
Is a second 16 GB card worth a 32 GB card?+
Can you load two models at the same time on 32 GB?+
Which quantization should you choose for 32 GB?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.