Advanced 11 minBy VRAM

Which LLM for 32 GB of VRAM ?

Direct response

With 32 GB of VRAM, you can fully load a dense 27B to 32B model in Q4 (16 to 20 GB) with a long context, a 30-35B MoE (19 to 21 GB), or gpt-oss-20b (14 GB) with plenty of headroom. The threshold that does not fit is 70B: 40 to 43 GB in Q4, so it necessarily spills partly into system RAM, where speed collapses. Always account for the context cache in addition to the weights.

32 GB of VRAM is the threshold where you stop choosing between model size and context length. This page explains what really fits, what changes depending on the card (RTX 5090 or Radeon AI PRO R9700), why a 70B or 123B does not become comfortable, and when a Mac’s unified memory or a second card is a better purchase.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#32 GB of VRAM: what the capacity changes compared with 24 GB

A Q4 model uses approximately 0.6 GB per billion parameters, plus the context cache, which grows with conversation length. At 24 GB, a 27–32B dense model in Q4 (16–20 GB) fits, but the context must remain short. At 32 GB, 12–16 GB remains for context, activations, and the system: this is the budget that makes 32,000- to 128,000-token windows possible on a 27B, or a 30–35B MoE with comfortable headroom.

32 GB budget for a few common models (Q4 weights, excluding context)
ModelQ4 weightRemaining capacity for context and systemVerdict
gpt-oss-20b14 GB18 GBVery wide margin, long context
Qwen 3.5 27B / Qwen 3.8 27B16 GB16 GBComfortable, long context possible
Gemma 4 31B18 GB14 GBComfortable
Qwen 3 32B19–20 GB12 GBGood, context to monitor
Qwen3-Coder 30B-A3B (MoE)19 GB13 GBGood, very fast for its size
Qwen 3.6 35B-A3B (MoE)21 GB11 GBGood, medium context
Llama 3.3 70B43 GBnegativeDoesn't fit: offloading required

These values are file weights, not total requirements: the site’s calculator adds the cache for a given model and context. The cache can also be quantized to q8_0, which, according to Ollama’s FAQ, cuts its size by about half compared with f16, provided Flash Attention is enabled.

#Which cards offer 32 GB, and why bandwidth matters

Two desktop cards stand out. The GeForce RTX 5090 includes 32 GB of GDDR7 on a 512-bit bus, with 1,792 GB/s of bandwidth according to NVIDIA. AMD's Radeon AI PRO R9700 also offers 32 GB, but in GDDR6 on a 256-bit bus at 640 GB/s, according to AMD, for desktop and professional use. Capacity is identical; generation speed is not.

Two 32 GB cards: same capacity, very different theoretical throughput
CardMemoryBandwidthCeiling for a 19 GB model
GeForce RTX 509032 GB GDDR7, 512-bit1,792 GB/sapproximately 94 t/s
Radeon AI PRO R970032 GB GDDR6, 256-bit640 GB/sapproximately 34 t/s

The ceiling is calculated by dividing bandwidth by the weights read for each token; it is never fully achieved, but the ratio remains valid: with the same model, the RTX 5090 generates more than twice as fast as the R9700. For an AMD card, the driver and software stack (ROCm, Vulkan) also matter; the dedicated guide to AMD cards details the limitations. Specific cards have their own pages: the RTX 5090 page provides the card's details.

Two other approaches reach 32 GB. Two 16 GB cards can split a model with llama.cpp, but the PCIe connection slows communication and the model must be split cleanly; the multi-GPU guide provides the method. A Mac with 32 GB or more of unified memory is a third option, with lower bandwidth but no dedicated VRAM limit.

Prices change quickly: every Monday and Thursday, the site’s tracker records the lowest price for graphics cards for local AI, along with the price per GB of VRAM.

#Models to prioritize on 32 GB

Dense 27B to 32B
The quality choice: Qwen 3.5 27B or 3.8 27B (16 GB in Q4), Gemma 4 31B (18 GB), Qwen 3 32B (19 to 20 GB). They leave room for a long context and work well with finer quantizations (Q5, Q6) if you want higher quality.
MoE 30–35B
Qwen3-Coder 30B-A3B (19 GB) and Qwen 3.6 35B-A3B (21 GB) activate only about 3 billion parameters per token: they offer the best speed for their size and are the right choice for code and agents.
gpt-oss-20b
14 GB in the Ollama library; the Ollama page says it can run on systems with at least 16 GB of memory. On 32 GB, it leaves room for a huge context and lets you keep a second model loaded at the same time.
24B code model
Devstral Small 2 24B (14 GB): a specialist in coding agents that leaves room for two resident models.

To choose, ask three questions. Does the model need to handle a long context (documents, code repository)? Choose a 27B dense model in Q4, leaving 16 GB for cache. Do you need speed (agents, tool loops)? Choose the MoE. Do you simply want more quality than at 24 GB? Move from Q4 to Q6 on a 27B to 32B model: there's room.

#Verify that the model really fits on the card

A model that “loads” is not necessarily entirely on the GPU. Ollama silently distributes layers between the card and RAM when VRAM is insufficient, and speed drops without an error message. The official FAQ states that the ollama ps command displays where the model was loaded in the Processor column: 100% GPU means it is entirely on the GPU, while a split such as 30%/70% CPU/GPU indicates overflow.

  1. 01
    Load the model with the target context
    Run the model with the actual context size of your use case, not the default value: the context cache is what causes overflow.
  2. 02
    Read the Processor column
    Run ollama ps in another terminal. Aim for 100% GPU; any CPU share cuts the speed.
  3. 03
    Reduce if necessary
    If the model runs out of room, reduce the context, enable q8_0 cache quantization with Flash Attention, or switch to a lighter quantization.
  4. 04
    Check remaining VRAM
    On Windows or Linux, nvidia-smi shows used memory; leave one to two GB of headroom for the display and spikes.
  5. 05
    Measure on your workload
    Compare throughput with a short prompt and with your real prompt: the second number is what matters.
!
The default context trap
According to its FAQ, Ollama uses a 4,096-token context window by default. A model may appear to fit in VRAM with this setting and run out of memory as soon as you increase it to 32,000 tokens to analyze a document. Always test with the context you need.

#70B and 123B: why offloading does not make these models usable

Llama 3.3 70B weighs 43 GB in the Ollama library, Mistral Large 123B 73 GB. On 32 GB, you therefore need to leave at least 11 GB for the 70B and 41 GB for the 123B in system memory. For each token, the processor reads these layers at system memory speed: about 102 GB/s at most for dual-channel DDR5-6400 (6,400 MT/s × 8 bytes × 2). Dividing by 41 GB gives the 123B a ceiling of about 2.5 tokens per second; for the 70B, with 11 GB in RAM, about 9 tokens per second. These are ceilings, not measurements: they mainly show that throughput does not degrade progressively, but in steps, as soon as the layers leave the card.

What happens when a model exceeds 32 GB (calculated from system memory)
ModelSize OllamaMoves to system RAMRAM ceiling (102 GB/s)
Llama 3.3 70B43 GBabout 11 GB (34%)about 9 t/s
Mistral Large 123B73 GBapproximately 41 GB (56%)about 2.5 t/s

This calculation assumes the card keeps 32 GB of weights; in practice, the context cache also takes up space, so more spills into RAM and the limit drops. MoEs behave better: llama.cpp offers an option to keep experts from certain layers on the CPU (--n-cpu-moe), for configurations that cannot load the entire model on the GPU. It saves a model that doesn't fit, but a llama.cpp repository user reports an approximately 80% speed loss when putting all experts on the CPU: it's a last-resort solution.

An oversized MoE: keep experts on the CPU (llama.cpp)
llama-server -m modele-moe.gguf -ngl 99 -fa --n-cpu-moe 12 -c 16384

The right value for --n-cpu-moe depends on your memory; you have to find it through testing. For a dense 70B, the right answer at 32 GB is almost always a smaller model or more memory, not offload.

#Fine-tuning a model on 32 GB: what works and what doesn't

QLoRA fine-tuning keeps the model weights in 4-bit form and trains only small adapters. The weights alone for a 70B model already take up 40 GB, so QLoRA on a 70B model does not fit in 32 GB. A 32B model in 4-bit form takes up about 19 to 20 GB: that leaves 12 GB for activations, the optimizer, and the adapters, which is feasible with a short context and a batch size of 1, with no safety margin. A 7B to 14B model is comfortable. These are rough estimates to confirm in the tool you use.

#32 GB of VRAM or another solution: the decision table

What to choose based on your needs
NeedSolutionWhy
Dense 27–32B, long context, maximum speed32 GB card NVIDIAHigh bandwidth, CUDA ecosystem
Same capacity at lower electrical powerRadeon AI PRO R9700Same 32 GB, lower theoretical bandwidth
64 GB and larger models, occasional useMac with 64 GB or more of unified memoryCapacity without offload, lower speed
Limited budget, 30B MoEA 24 GB card is often enoughThe 19 GB MoE fits with a moderate context

#Frequently asked questions

FAQ
What is the best LLM for 32 GB of VRAM?+
For quality, a dense 27B to 32B model in Q4 (Qwen 3.5 or 3.8 27B, Gemma 4 31B, Qwen 3 32B) with a long context. For speed and coding, a 30-35B MoE such as Qwen3-Coder 30B-A3B. No model is universally better: the task and language determine the choice.
Is 32 GB of VRAM enough for Mistral Large 123B?+
No. The Ollama library lists 73 GB for Mistral Large 123B. More than 40 GB should remain in RAM, where read speeds are limited to about 100 GB/s, giving a theoretical ceiling of 2 to 3 tokens per second. It's usable for background processing, not for chat.
RTX 5090 or Mac Studio 64 GB for local AI?+
The card for speed: 1,792 GB/s versus a few hundred GB/s for a Mac, on models that fit in 32 GB. The Mac for capacity: a 40- to 50-GB model fits without offloading, at a lower speed. Choose based on the size of the models you load.
Is a second 16 GB card worth a 32 GB card?+
Two 16 GB cards add up to 32 GB, but the model is split between them and the PCIe link slows communication; splitting also requires configuration in llama.cpp. A 32 GB card is simpler and faster for the same total, unless you already own one of the two cards.
Can you load two models at the same time on 32 GB?+
Yes, for example gpt-oss-20b (14 GB) and Devstral Small 2 24B (14 GB), for a total of 28 GB of weights, with a short context. Ollama keeps models in memory for five minutes by default and can keep several if memory allows; monitor the context cache, which is added to each one.
Which quantization should you choose for 32 GB?+
Q4_K_M remains the best compromise for dense models from 27B to 32B. If you have headroom (a model with 14 to 16 GB of weights), move up to Q5 or Q6 for better quality. Go below Q4 only to fit an oversized model, accepting a loss in quality.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.