AI PC configuration: 32 GB budget VRAM
With 32 GB of VRAM, you can load a dense 32-billion-parameter model in Q4 (about 19 to 20 GB) or Qwen3-Coder 30B-A3B (19 GB), but not a 30-billion-parameter model in Q8 (about 32 GB) or a 70B in Q4 (about 40 GB). You can reach 32 GB with a 32 GB card, or 48 GB with two 24 GB cards.
This page explains how much VRAM to target based on the model, compares a single 32 GB card, two 24 GB cards, and a 24 GB card, sizes the power supply, RAM, storage, and case, and corrects a misconception: a 30B model in Q8 does not fit in 32 GB.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What 32 GB of VRAM really allow
With 32 GB of VRAM, you can comfortably load a dense 32-billion-parameter model in Q4 (about 19 to 20 GB of weights) with several thousand tokens of context, or a 30-billion-parameter mixture-of-experts model such as Qwen3-Coder 30B-A3B, which the Ollama library lists at 19 GB. You cannot load a 30-billion-parameter model in Q8: it weighs about 32 GB, filling the entire card with no room for context or memory reserved by the driver. A 70B in Q4, at about 40 GB, does not fit either. You can get 32 GB with a single 32 GB card or two 24 GB cards—two options with very different trade-offs.
| Model | Weights | 24 GB (1 card) | 32 GB (1 card) | 48 GB (2 × 24 GB) |
|---|---|---|---|---|
| Qwen3-Coder 30B-A3B in Q4 | ≈ 19 GB | Yes, medium context | Yes, long context | Yes |
| Dense 32B model in Q4 | ≈ 19 to 20 GB | Fair, short context | Yes | Yes |
| Dense 32B model in Q8 | ≈ 34 GB | No | No | Yes, medium context |
| 30B model in Q8 | ≈ 32 GB | No | No, there’s no headroom | Yes |
| 70B in Q4 | ≈ 40 GB | No | No | Yes, short context |
Q8 weighs about one byte per parameter, which is why a 30-billion-parameter model does not fit in 32 GB. Since the card cannot dedicate all its memory to the model, target weights at least 3 to 4 GB below the VRAM capacity to leave room for the KV cache. The site's VRAM calculator performs this calculation for a given model and context.
#The GPU: three ways to reach 24 to 48 GB
#Option A: one 32 GB card
The RTX 5090 comes with 32 GB of GDDR7 on a 512-bit interface, according to NVIDIA. Its total graphics power reaches 575 W, and NVIDIA lists 1,000 W as the required system power. It is powered by four 8-pin PCIe cables through the included adapter, or by a 600 W PCIe Gen 5 cable. A single GPU avoids any multi-card configuration: it is the simplest option and the one that delivers the best speed per model, because memory is not split across cards.
#Path B: two 24 GB cards
The RTX 3090 provides 24 GB of GDDR6X; NVIDIA lists a board power of 350 W and a required system power of 750 W for a single card. Two cards provide 48 GB in total, with 700 W for the cards alone. In llama.cpp, the distribution option is --split-mode: by default, layer mode splits the layers across the GPUs, while --tensor-split sets the proportion assigned to each card. Tensor mode is described there as experimental. Distribution does not speed up generation the way a single faster card would; its primary purpose is to fit a larger model.
#Path C: a single 24 GB card, tight budget
A used RTX 3090 or 4090, with 24 GB, already handles a 30-billion-parameter mixture-of-experts model in Q4 and a dense 27- to 32-billion-parameter model with a short context. This is the route for a budget capped at around a few thousand euros: prices change every week, so the /prix-gpu-ia tracker is the only reliable reference. Headroom shrinks quickly with context: a KV cache of several gigabytes comes out of the 4 to 5 GB remaining.
| Criterion | A 32 GB card | Two 24 GB cards | A 24 GB card |
|---|---|---|---|
| Usable VRAM | 32 GB | 48 GB allocated | 24 GB |
| Complexity | None | Distribution to configure, two large cards | None |
| Card performance | 575 W (RTX 5090) | 2 × 350 W (RTX 3090) | 350 W (RTX 3090) |
| Power supply | At least 1,000 W (NVIDIA) | At least 1,200 W recommended (calculated for this guide) | 750 W (NVIDIA) |
| Largest model in Q4 | 32B with context | 70B with a short context | 30B experts, 32B short |
#The context: what VRAM leaves for the KV cache
Model weights don't tell the whole story: each context token adds an entry to the KV cache, which is also stored in VRAM. The Qwen3-Coder-30B-A3B specifications list 48 layers and grouped attention with 32 query heads for 4 KV heads. This structure divides the KV cache size by eight compared with a model having as many KV heads as query heads, which explains how such a model supports very long contexts on a 24 or 32 GB card. Not all models are as memory-efficient: the more KV heads a model has, the faster its cache grows with the context.
The rule of thumb remains the same: start with the model's weight, add 3 to 4 GB of headroom, then see what remains for the target context. If you run short, quantize the KV cache rather than reducing the model's quantization: the dedicated guide explains the trade-off.
#CPU, RAM, and motherboard: don't waste the budget
During GPU generation, the processor does little work: it loads the models, tokenizes, prepares the context, and runs the system. A mid-range processor is sufficient, and the money is better spent on VRAM. The motherboard matters more: for two 24 GB cards, you need two physically spaced PCIe x16 slots and enough room for thick cards (NVIDIA describes the RTX 3090 Founders Edition as a three-slot card). The number of PCIe lanes per slot has little effect on generation once the model is loaded.
For system memory, 32 GB is the baseline for a PC equipped with a 32 GB card; 64 GB is more comfortable if you keep an IDE, browser, and container open while the model runs, or if you offload part of the model to the processor. In the latter case, speed depends on RAM: see the guide to running an LLM without a GPU. A common query, “LLM 32 GB RAM,” often refers to this situation: with 32 GB of RAM and no graphics card, 7- to 14-billion-parameter models in Q4 remain usable; a 32B in Q4 (about 19 to 20 GB) fits but generates slowly.
#Power supply: the component not to skimp on
NVIDIA indicates that 1,000 W of system power is required for an RTX 5090 and 750 W for an RTX 3090. For two RTX 3090, the calculation is simple: 2 × 350 W for the cards, plus about 250 W for the processor, motherboard, drives, and fans, gives about 950 W under load, with spikes above that. A 1,200 W or higher power supply leaves healthy headroom; this is a recommendation from this guide, not an Apple or NVIDIA specification. Also check the number of 8-pin PCIe connectors: an RTX 5090 requires four through the adapter, or a 600 W PCIe Gen 5 cable.
#Storage: loading speed, not capacity
A 20 to 40 GB model is loaded each time you switch models. Calculation using assumptions to adapt: at 200 MB/s, the typical hard-drive speed, 40 GB takes approximately 200 seconds; at 3.5 GB/s, the speed of an NVMe SSD, approximately 12 seconds. A 2 TB NVMe SSD is enough for around ten models and a RAG index; a second drive is used for backups. A PCIe 5.0 SSD provides no benefit for generation: generation takes place in VRAM.
#Case and cooling
A 575 W card releases as much heat as a small space heater: the case must exhaust hot air faster than it enters, with front intake and rear and top fans. Two thick cards side by side heat each other up: leave one slot between them, or choose thinner models. The guide to inference thermals details the temperatures to monitor and the fan curves.
- Which LLM for 32 GB of VRAM
- Which LLM for 24 GB of VRAM
- Which LLM on RTX 5090
- Which LLM on RTX 3090 and a 3090 Ti
- Multi-GPU LLM with llama.cpp
- Local AI graphics card prices
#Validate your configuration before buying
- 01Start with the model, not the cardChoose the target model and quantization, then add 3 to 4 GB for the KV cache and driver: the total gives you the minimum VRAM. The site's VRAM calculator performs this calculation.
- 02Choose the GPU pathOne 32 GB card for simplicity, two 24 GB cards if budget is the priority and the case and motherboard support it, or a single 24 GB card for a tight budget.
- 03Size the power supplyAt a minimum, use the system power NVIDIA specifies for the card, and add headroom for two cards. Count the PCIe connectors.
- 04Check the caseLength, thickness, and airflow requirements for each card. Two thick cards require a wide case and a motherboard with spaced-apart slots.
- 05After assembly, check the loadThe nvidia-smi command displays the VRAM usage and power draw of each card; ollama ps shows whether the model is running entirely on the GPU.
- Source: NVIDIA, GeForce RTX 5090 spec sheet
- Source: NVIDIA, the GeForce RTX 3090 and 3090 Ti specifications
- Source: Qwen3-Coder-30B-A3B’s Hugging Face page
- Source: Ollama library, Qwen3-Coder
#Frequently asked questions
Is 32 GB of VRAM enough for a 32-billion-parameter model?+
RTX 5090 or two RTX 3090 for 32 GB of VRAM?+
What configuration do you need to run Qwen3-Coder 30B-A3B?+
Is 32 GB of RAM enough for an LLM without a graphics card?+
Which power supply should you choose for a RTX 5090?+
Where can you find current prices for AI graphics cards?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.