Which LLM for 12 GB of VRAM ?
With 12 GB of VRAM, choose a 12- to 14-billion-parameter model in Q4 (Gemma 4 12B: 7 GB, Qwen 3 14B: 9 GB) or a 9B in Q8 (10 GB), with 1 to 3 GB of headroom for the context. Models with 20 to 24 billion parameters require 16 GB; those with 27 to 35 billion require 24 GB.
Twelve gigabytes of VRAM comfortably opens the door to models with 12 to 14 billion parameters. This guide shows what fits, which cards actually have 12 GB, how to allocate memory for local RAG, and when to move up to 16 GB.
Choosing a machine? Our picks by budget →
To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#12 GB of VRAM: the tier where 14 billion parameters become possible
With 12 GB of VRAM, you can run 12- to 14-billion-parameter models in Q4 with plenty of room for context, or a 9B model in Q8. According to the QuelLLM catalog, a Gemma 4 12B in Q4 weighs 7 GB, a Qwen 3 14B in Q4 weighs 9 GB, and a Qwen 3.5 9B in Q8 weighs 10 GB. The 12 GB tier is where you move beyond 8- to 9-billion-parameter models, which are more limited in reasoning and writing quality, without yet reaching the 24-billion-parameter models that require 16 GB. Still out of reach: 27- to 35-billion-parameter models, which require between 16 and 21 GB in Q4.
| Model | Q4 | Q5 | Q8 | At 12 GB |
|---|---|---|---|---|
| Qwen 3.5 9B | 6 GB | 7 GB | 10 GB | Q8 possible, short context |
| Gemma 4 12B | 7 GB | 9 GB | 13 GB | Q4 or Q5; Q8 spills over |
| Mistral Nemo 12B | 7 GB | 9 GB | 13 GB | Q4 or Q5; Q8 spills over |
| Qwen 3 14B | 9 GB | 11 GB | 16 GB | Q4 comfortable, Q5 tight |
| Mistral Small 24B | 14 GB | 17 GB | 26 GB | No |
How to read this table: add the KV cache and about one gigabyte of headroom to the model weights. A Qwen 3 14B in Q4 (9 GB) therefore leaves about 2 GB for context; a Qwen 3.5 9B in Q8 (10 GB) leaves only one, which limits context. Before going further, one correction: an earlier version of this guide described the 9B in Q8 on 12 GB as comfortable; the catalog lists 10 GB of weights, so the headroom is tight.
#12 GB cards: the ones that really have it
Many card families have versions with different capacities; the name alone is not enough. NVIDIA describes, for its RTX 4070 family, the 4070 Ti Super with 16 GB and the 4070 Ti, 4070 Super, and 4070 with 12 GB. The RTX 5070 offers 12 GB of GDDR7 on a 192-bit bus, while the 5070 Ti has 16. The RTX 3060 comes in 12 GB and 8 GB versions, with the 12 GB version being more interesting for local AI. On the AMD side, the RX 7700 XT guide places it in this tier.
| Card | Capacity | Guide |
|---|---|---|
| RTX 3060 (12 GB version) | 12 GB GDDR6 | LLM on RTX 3060 12 GB |
| RTX 4070, 4070 Super, 4070 Ti | 12 GB GDDR6X or GDDR6 | LLM on RTX 4070 |
| RTX 5070 | 12 GB GDDR7, 192-bit | LLM on RTX 5070 |
| RX 7700 XT | 12 GB (see the guide) | LLM on RX 7700 XT |
For comparison, capacity is the primary criterion and bandwidth the second: at equal capacity, faster memory means smoother generation. Prices change every week; the site’s price tracker shows the cost per gigabyte of VRAM for 12 GB and 16 GB cards, letting you determine whether the next tier is worth the extra cost.
#Which models to choose on 12 GB based on your use case
| Usage | Model and quantization | Weights | What you need to know |
|---|---|---|---|
| General discussion in French | Mistral Nemo 12B Q5 | 9 GB | Labeled French in the catalog; 3 GB of headroom |
| Reasoning and writing | Qwen 3 14B Q4 | 9 GB | About 2 GB for the context |
| Multimodal (images) | Gemma 4 12B Q4 | 7 GB | The image module uses additional memory |
| Maximum quality from a 9B model | Qwen 3.5 9B Q8 | 10 GB | Context limited to a few thousand tokens |
| Speed and lightness | Granite 4.2 8B Q5 | 6 GB | Leave 5 GB for an embedder or a long context |
For coding, the most capable specialized models in the catalog exceed 12 GB: Devstral Small 2 24B weighs 14 GB in Q4 and Qwen3-Coder 30B-A3B 19 GB. On 12 GB, stick with a general-purpose model of 12 to 14 billion parameters or Qwen 2.5 Coder 14B, whose catalog listing gives 9 GB in Q4. These models are comparable to the 8 and 16 GB options in the neighboring guides.
#Speed: the calculation you need to do yourself
Whether you have 12 GB or more, generation speed is capped by the card’s memory bandwidth divided by the size of the weights read for each token. Each manufacturer publishes this figure on the card’s specifications page, in Go/s. Simply divide it by the model size to get a theoretical ceiling, which real-world speed never fully reaches. The table gives this ceiling for 100 GB/s of bandwidth; multiply it by your card’s figure.
| Q4 model | Weights read | Ceiling per 100 GB/s | Example for 300 GB/s |
|---|---|---|---|
| Granite 4.2 8B | 4.6 GB | ≈ 22 tokens/s | ≈ 65 tokens/s |
| Gemma 4 12B | 7 GB | ≈ 14 tokens/s | ≈ 43 tokens/s |
| Qwen 3 14B | 9 GB | ≈ 11 tokens/s | ≈ 33 tokens/s |
| Qwen 3.5 9B in Q8 | 10 GB | 10 tokens/s | 30 tokens/s |
This ceiling explains why finer quantization is slower: going from Q4 to Q8 nearly doubles the weights read. It also explains why two cards with the same capacity can differ: the memory bus matters. NVIDIA specifies 192 bits for the RTX 5070 with 12 GB and 256 bits for the 16 GB 5070 Ti. We do not publish tokens per second by card because we lack a comparable source; for your hardware, ollama run with the --verbose option displays the actual speed.
#Verify that a model fits before adopting it
- 01Read the catalog weightsOpen the model page in the QuelLLM catalog and note the size of the target quantization.
- 02Add headroomAdd about 1 GB for the system and 1 to 3 GB for the KV cache, depending on the target context. The total must stay under 12 GB.
- 03Run and measureLoad the model, ask a question, then run ollama ps: the PROCESSOR column should show 100% GPU.
- 04Adjust if necessaryReduce the context, quantize the KV cache, or drop down one quantization level, in that order.
#Local RAG on 12 GB: how to allocate memory
A local RAG system combines three components: an embedding model that indexes your documents, optionally a reranker that reorders passages, and the LLM that writes the response. The first two are much smaller than the LLM and can run on either the GPU or the processor. Calculate the budget by adding all three, then leave room for the KV cache: the LLM's context grows with the number of passages you send it.
| Strategy | Placement | Advantage | Limit |
|---|---|---|---|
| Everything about GPUs | 8B LLM in Q5 (6 GB) + embedder + reranker | Fast responses | Little headroom for a long context |
| Embedder on the processor | 12B LLM in Q4 (7 GB) on GPU | Frees up VRAM for context | Slower indexing |
| Index, then shut down | Embedder launched only during indexing | VRAM completely free afterward | Two steps to manage |
The second strategy is often best: indexing is a background task that doesn't need a GPU, while response generation benefits fully from VRAM. An embedder runs in the background during indexing, then goes idle; on a CPU, indexing a document folder takes longer, but it only happens once. We don't provide sizes for embedders here; consult the guides on embeddings and local RAG for each model's values.
#What 12 GB don't allow
- 24-billion-parameter models
- Mistral Small 24B weighs 14 GB in Q4 according to the catalog: it exceeds the limit. For 24 billion parameters, you need 16 GB.
- 27- to 32-billion-parameter dense models
- Qwen 3.8 27B and similar models weigh 16 GB in Q4; target 24 GB.
- 30- to 35-billion-parameter MoE models
- Qwen 3.6 35B-A3B weighs 21 GB, GLM 4.7 Flash 19 GB, and Qwen3-Coder 30B-A3B 19 GB: out of reach on 12 GB, even though they read only 3 billion active parameters, because the entire model must reside in memory.
- Long context on a 14B
- A Qwen 3 14B in Q4 leaves about 2 GB: don't expect tens of thousands of tokens without quantizing the KV cache.
Ollama lets you quantize the KV cache to save space: according to its documentation, OLLAMA_KV_CACHE_TYPE=q8_0 uses about half the memory of the default f16, provided Flash Attention is enabled. The KV cache guide details this setting, and Ollama’s default context window, 4 096 tokens, can be changed with OLLAMA_CONTEXT_LENGTH.
#12 GB or 16 GB: the right choice for you
| Your needs | Is 12 GB enough? | Why |
|---|---|---|
| Chat, translation, and summarization with a 9B to 14B model | Yes | Q4 or Q5, with context |
| Lightweight RAG with an on-CPU embedder | Yes | 7 to 9 GB LLM on GPU |
| 20- to 24-billion-parameter models (gpt-oss 20B, Mistral Small 24B) | No | 13 to 14 GB of Q4 weights |
| Coding agents with long context | No | Context and weights exceed 12 GB |
| Measured budget, occasional use | Yes | The extra capacity adds nothing if you don't need 20- to 24-billion-parameter models |
- Which LLM for 16 GB of VRAM
- Which LLM for 8 GB of VRAM?
- LLM on RTX 3060 12 GB
- Quantize the KV cache to save VRAM
- Local RAG: introduction
- Local AI graphics card prices
- Source: Ollama FAQ (context, KV cache)
- Source: NVIDIA spec sheet for the RTX 4070 family
- Source: NVIDIA page for the RTX 5070 family
- QuelLLM model catalog (VRAM sizes)
#Frequently asked questions
What is the best LLM for 12 GB of VRAM?+
Is 12 GB of VRAM enough in 2026?+
Can you run a 14B on 12 GB?+
12 GB or 16 GB for LLMs?+
Which 12 GB card should you choose?+
Can 30-billion-parameter MoE models fit in 12 GB?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.