Beginner 11 minBy VRAM

Which LLM for 12 GB of VRAM ?

Direct response

With 12 GB of VRAM, choose a 12- to 14-billion-parameter model in Q4 (Gemma 4 12B: 7 GB, Qwen 3 14B: 9 GB) or a 9B in Q8 (10 GB), with 1 to 3 GB of headroom for the context. Models with 20 to 24 billion parameters require 16 GB; those with 27 to 35 billion require 24 GB.

Twelve gigabytes of VRAM comfortably opens the door to models with 12 to 14 billion parameters. This guide shows what fits, which cards actually have 12 GB, how to allocate memory for local RAG, and when to move up to 16 GB.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#12 GB of VRAM: the tier where 14 billion parameters become possible

With 12 GB of VRAM, you can run 12- to 14-billion-parameter models in Q4 with plenty of room for context, or a 9B model in Q8. According to the QuelLLM catalog, a Gemma 4 12B in Q4 weighs 7 GB, a Qwen 3 14B in Q4 weighs 9 GB, and a Qwen 3.5 9B in Q8 weighs 10 GB. The 12 GB tier is where you move beyond 8- to 9-billion-parameter models, which are more limited in reasoning and writing quality, without yet reaching the 24-billion-parameter models that require 16 GB. Still out of reach: 27- to 35-billion-parameter models, which require between 16 and 21 GB in Q4.

What 12 GB can fit (QuelLLM catalog weights)
ModelQ4Q5Q8At 12 GB
Qwen 3.5 9B6 GB7 GB10 GBQ8 possible, short context
Gemma 4 12B7 GB9 GB13 GBQ4 or Q5; Q8 spills over
Mistral Nemo 12B7 GB9 GB13 GBQ4 or Q5; Q8 spills over
Qwen 3 14B9 GB11 GB16 GBQ4 comfortable, Q5 tight
Mistral Small 24B14 GB17 GB26 GBNo

How to read this table: add the KV cache and about one gigabyte of headroom to the model weights. A Qwen 3 14B in Q4 (9 GB) therefore leaves about 2 GB for context; a Qwen 3.5 9B in Q8 (10 GB) leaves only one, which limits context. Before going further, one correction: an earlier version of this guide described the 9B in Q8 on 12 GB as comfortable; the catalog lists 10 GB of weights, so the headroom is tight.

#12 GB cards: the ones that really have it

Many card families have versions with different capacities; the name alone is not enough. NVIDIA describes, for its RTX 4070 family, the 4070 Ti Super with 16 GB and the 4070 Ti, 4070 Super, and 4070 with 12 GB. The RTX 5070 offers 12 GB of GDDR7 on a 192-bit bus, while the 5070 Ti has 16. The RTX 3060 comes in 12 GB and 8 GB versions, with the 12 GB version being more interesting for local AI. On the AMD side, the RX 7700 XT guide places it in this tier.

12 GB cards and dedicated guides
CardCapacityGuide
RTX 3060 (12 GB version)12 GB GDDR6LLM on RTX 3060 12 GB
RTX 4070, 4070 Super, 4070 Ti12 GB GDDR6X or GDDR6LLM on RTX 4070
RTX 507012 GB GDDR7, 192-bitLLM on RTX 5070
RX 7700 XT12 GB (see the guide)LLM on RX 7700 XT

For comparison, capacity is the primary criterion and bandwidth the second: at equal capacity, faster memory means smoother generation. Prices change every week; the site’s price tracker shows the cost per gigabyte of VRAM for 12 GB and 16 GB cards, letting you determine whether the next tier is worth the extra cost.

#Which models to choose on 12 GB based on your use case

A starting point for each use case
UsageModel and quantizationWeightsWhat you need to know
General discussion in FrenchMistral Nemo 12B Q59 GBLabeled French in the catalog; 3 GB of headroom
Reasoning and writingQwen 3 14B Q49 GBAbout 2 GB for the context
Multimodal (images)Gemma 4 12B Q47 GBThe image module uses additional memory
Maximum quality from a 9B modelQwen 3.5 9B Q810 GBContext limited to a few thousand tokens
Speed and lightnessGranite 4.2 8B Q56 GBLeave 5 GB for an embedder or a long context

For coding, the most capable specialized models in the catalog exceed 12 GB: Devstral Small 2 24B weighs 14 GB in Q4 and Qwen3-Coder 30B-A3B 19 GB. On 12 GB, stick with a general-purpose model of 12 to 14 billion parameters or Qwen 2.5 Coder 14B, whose catalog listing gives 9 GB in Q4. These models are comparable to the 8 and 16 GB options in the neighboring guides.

→
A simple criterion for choosing quantization
Keep at least 2 GB free beyond the model weights if you use a context of more than 8 000 tokens. If the margin is smaller, drop one quantization level (Q8 to Q5, Q5 to Q4) instead of reducing the context: Q5’s quality loss is small for typical use.

#Speed: the calculation you need to do yourself

Whether you have 12 GB or more, generation speed is capped by the card’s memory bandwidth divided by the size of the weights read for each token. Each manufacturer publishes this figure on the card’s specifications page, in Go/s. Simply divide it by the model size to get a theoretical ceiling, which real-world speed never fully reaches. The table gives this ceiling for 100 GB/s of bandwidth; multiply it by your card’s figure.

Theoretical ceiling per 100 GB/s of memory bandwidth
Q4 modelWeights readCeiling per 100 GB/sExample for 300 GB/s
Granite 4.2 8B4.6 GB≈ 22 tokens/s≈ 65 tokens/s
Gemma 4 12B7 GB≈ 14 tokens/s≈ 43 tokens/s
Qwen 3 14B9 GB≈ 11 tokens/s≈ 33 tokens/s
Qwen 3.5 9B in Q810 GB10 tokens/s30 tokens/s

This ceiling explains why finer quantization is slower: going from Q4 to Q8 nearly doubles the weights read. It also explains why two cards with the same capacity can differ: the memory bus matters. NVIDIA specifies 192 bits for the RTX 5070 with 12 GB and 256 bits for the 16 GB 5070 Ti. We do not publish tokens per second by card because we lack a comparable source; for your hardware, ollama run with the --verbose option displays the actual speed.

#Verify that a model fits before adopting it

  1. 01
    Read the catalog weights
    Open the model page in the QuelLLM catalog and note the size of the target quantization.
  2. 02
    Add headroom
    Add about 1 GB for the system and 1 to 3 GB for the KV cache, depending on the target context. The total must stay under 12 GB.
  3. 03
    Run and measure
    Load the model, ask a question, then run ollama ps: the PROCESSOR column should show 100% GPU.
  4. 04
    Adjust if necessary
    Reduce the context, quantize the KV cache, or drop down one quantization level, in that order.

#Local RAG on 12 GB: how to allocate memory

A local RAG system combines three components: an embedding model that indexes your documents, optionally a reranker that reorders passages, and the LLM that writes the response. The first two are much smaller than the LLM and can run on either the GPU or the processor. Calculate the budget by adding all three, then leave room for the KV cache: the LLM's context grows with the number of passages you send it.

Three possible distributions
StrategyPlacementAdvantageLimit
Everything about GPUs8B LLM in Q5 (6 GB) + embedder + rerankerFast responsesLittle headroom for a long context
Embedder on the processor12B LLM in Q4 (7 GB) on GPUFrees up VRAM for contextSlower indexing
Index, then shut downEmbedder launched only during indexingVRAM completely free afterwardTwo steps to manage

The second strategy is often best: indexing is a background task that doesn't need a GPU, while response generation benefits fully from VRAM. An embedder runs in the background during indexing, then goes idle; on a CPU, indexing a document folder takes longer, but it only happens once. We don't provide sizes for embedders here; consult the guides on embeddings and local RAG for each model's values.

#What 12 GB don't allow

24-billion-parameter models
Mistral Small 24B weighs 14 GB in Q4 according to the catalog: it exceeds the limit. For 24 billion parameters, you need 16 GB.
27- to 32-billion-parameter dense models
Qwen 3.8 27B and similar models weigh 16 GB in Q4; target 24 GB.
30- to 35-billion-parameter MoE models
Qwen 3.6 35B-A3B weighs 21 GB, GLM 4.7 Flash 19 GB, and Qwen3-Coder 30B-A3B 19 GB: out of reach on 12 GB, even though they read only 3 billion active parameters, because the entire model must reside in memory.
Long context on a 14B
A Qwen 3 14B in Q4 leaves about 2 GB: don't expect tens of thousands of tokens without quantizing the KV cache.

Ollama lets you quantize the KV cache to save space: according to its documentation, OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16, provided Flash Attention is enabled. The KV cache guide details this setting, and Ollama’s default context window, 4 096 tokens, can be changed with OLLAMA_CONTEXT_LENGTH.

16k context and q8_0 KV cache for a 14B
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=16384 ollama serve
ollama ps

#12 GB or 16 GB: the right choice for you

Choosing between 12 and 16 GB
Your needsIs 12 GB enough?Why
Chat, translation, and summarization with a 9B to 14B modelYesQ4 or Q5, with context
Lightweight RAG with an on-CPU embedderYes7 to 9 GB LLM on GPU
20- to 24-billion-parameter models (gpt-oss 20B, Mistral Small 24B)No13 to 14 GB of Q4 weights
Coding agents with long contextNoContext and weights exceed 12 GB
Measured budget, occasional useYesThe extra capacity adds nothing if you don't need 20- to 24-billion-parameter models

#Frequently asked questions

Frequently asked questions
What is the best LLM for 12 GB of VRAM?+
A 12- to 14-billion-parameter model in Q4 or Q5, such as Gemma 4 12B (7 GB in Q4), Mistral Nemo 12B, or Qwen 3 14B (9 GB), according to the QuelLLM catalog. They leave room for context. A Qwen 3.5 9B in Q8 (10 GB) provides the best accuracy for a 9B, with a shorter context.
Is 12 GB of VRAM enough in 2026?+
Yes for chat, translation, summarization, and light RAG with models up to 14 billion parameters. No for 20- to 24-billion-parameter models, which weigh 13 to 14 GB in Q4, or for 27- to 35-billion-parameter models, weighing 16 to 21 GB. The latter require 16 to 24 GB.
Can you run a 14B on 12 GB?+
Yes, in Q4: Qwen 3 14B weighs 9 GB according to the catalog, leaving about 2 GB for the KV cache and system. Q5 (11 GB) is too tight, and Q8 (16 GB) is impossible. Check with ollama ps that the model is 100% on the GPU.
12 GB or 16 GB for LLMs?+
Choose 16 GB if you want 20- to 24-billion-parameter models, such as gpt-oss 20B (13 GB) or Mistral Small 24B (14 GB), with some context. Stay at 12 GB if you are content with 9- to 14-billion-parameter models. Compare the price per gigabyte in the site's tracking data.
Which 12 GB card should you choose?+
Check capacity first: the RTX 3060 comes in 12 and 8 GB versions, and the 4070 Ti Super has 16 GB. Among the 12 GB cards, the RTX 3060 is the most affordable, while the RTX 4070 and 5070 are the fastest. Today's price is shown in the site's price tracker.
Can 30-billion-parameter MoE models fit in 12 GB?+
No. An MoE model reads only part of its weights for each token, but it has to be loaded in full: Qwen3-Coder 30B-A3B weighs 19 GB and Qwen 3.6 35B-A3B 21 GB in Q4. With 12 GB, part of it would spill over to the CPU and speed would drop sharply.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.