Which LLM for 6 GB of VRAM ?
With 6 GB of VRAM, a 4-billion-parameter model in Q4 or Q5 fits comfortably with a context of several thousand tokens. A 7–8B model in Q4 (4.6 to 5.2 GB) still loads, but with no headroom: the context and memory reserved by the driver push it onto the CPU. Another surprising point: on 6 GB cards, memory bandwidth—not capacity—determines speed, and it varies from half to double.
6 GB is the threshold for GTX 1660, RTX 2060, and RTX 3050 6 GB. It's enough for a real local assistant, provided you choose the right model and account for memory that isn't used by the model. This page explains which models fit, which 6 GB card is faster than another, how to verify that a model stays on the card, and what 12 GB would change.
Choosing a machine? Our picks by budget →
To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What 6 GB makes possible: the complete memory calculation
VRAM holds more than just the weights. It also contains the context cache, which grows with the conversation, the engine's computation buffers, and, under Windows, part of the display and browser workload. On a 6 GB card, you can budget approximately 5 to 5.5 GB for the model and its cache, with the rest taken by the system. The site's rule of thumb for Q4 weights is 3B ≈ 2 GB, 7-8B ≈ 5 GB.
| Model | Q4 weight | Headroom above 5.5 GB | Verdict |
|---|---|---|---|
| Gemma 4 2B | 1.2 GB | 4.3 GB | Very large headroom, short and fast responses |
| Qwen 3.5 4B | 2.3 GB | 3.2 GB | Comfortable, with context of several thousand tokens |
| Phi-4 Mini 3.8B | 3 GB | 2.5 GB | Comfortable |
| Granite 4.2 8B | 4.6 GB | 0.9 GB | Simply: very short context |
| Qwen 3 8B (Ollama) | 5.2 GB | 0.3 GB | Overflows as soon as the context gets longer |
| Qwen 3.5 9B | 6 GB | negative | Doesn't fit |
The sizes come from the QuelLLM catalog and the Ollama library. The “headroom” column is an estimate: it assumes 5.5 GB available, which depends on your system and applications. An 8B model is therefore possible on paper, but it leaves you with less than one gigabyte for the cache: the context setting determines whether the experience is smooth or the layers spill over onto the processor.
#Three 6 GB cards, three speeds: bandwidth
Generation speed is limited by memory bandwidth divided by the size of the weights read for each token. So two 6 GB cards are not equivalent. According to Wikipedia’s GeForce 20 generation table, the RTX 2060 6 GB delivers 336 GB/s on a 192-bit bus. The newer RTX 3050 6 GB has only a 96-bit bus (NVIDIA) and 168 GB/s according to the same encyclopedia: half as much.
| 6 GB card | Bandwidth | 2.3 GB model (4B Q4) | 4.6 GB model (8B Q4) |
|---|---|---|---|
| RTX 2060 6 GB | 336 GB/s | approximately 146 t/s | about 73 t/s |
| RTX 3050 6 GB | 168 GB/s | about 73 t/s | about 37 t/s |
These ceilings are never reached, and we have no measurements to cite for these cards: the actual percentage depends on the engine, context, and card. But the 2-to-1 ratio holds. For local AI, a used RTX 2060 is therefore faster than a new 6 GB RTX 3050 despite its age, and the newer card has only energy efficiency and newer features in its favor. For GTX 1660, the variant matters: check the exact model number, because memory and bus width vary by model.
#The right model for every use case on 6 GB
- Chat, questions, writing
- A 3- to 4B model such as Qwen 3.5 4B or Phi-4 Mini: accurate, fast responses and comfortable context. For better French quality, try the 8B in Q4 with a reduced context.
- Code
- A small 3B to 7B coding model for autocompletion, and a general-purpose 4B model for explanations. Don’t expect a reliable coding agent at this size.
- RAG and documents
- A 4B model and a lightweight embedding model: everything fits in memory. Context is the limiting factor, so send short excerpts.
- Vision and audio
- Lightweight multimodal models exist, but the image encoder consumes memory in addition to the model: test with a 2 to 4B model.
A common trap: MoE models such as Qwen3-Coder 30B-A3B activate only 3 billion parameters, which may make you think they would fit on 6 GB. They do not: all experts must be in memory, and the file weighs 19 GB. The only way to use them is to keep the experts on the CPU, which requires a lot of RAM and comes at a steep speed cost. The guide to MoE models explains the principle.
#The hybrid GPU and CPU mode: useful, but not magic
When a model exceeds the card's capacity, Ollama and llama.cpp can place some layers on the CPU. This allows loading a 9B on 6 GB, but speed drops in stages: layers left in system RAM are read at system memory speed, a fraction of the card's speed. For MoE models, llama.cpp provides an option, --n-cpu-moe, which keeps the experts from certain layers on the CPU and leaves on the card what usually weighs the most. A user of the llama.cpp repository reports about 20 tokens per second with gpt-oss-20b and all experts on the CPU, on a machine with dual-channel DDR4-3200, while freeing most of the VRAM.
This account shows that the approach works, not that it suits your machine. It requires 32 GB of RAM for a model this size, and its throughput depends on memory bandwidth. For everyday use on 6 GB, a 4B model running entirely on the card remains simpler and faster.
#Five settings that matter on 6 GB
- 01Limit contextOllama's FAQ states that the default window is 4,096 tokens and can be adjusted. With 6 GB, increase it in increments (6,000, 8,000) and monitor memory: each increment adds cache.
- 02Quantize the K/V cacheWith Flash Attention enabled, OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache memory usage compared with f16, according to the Ollama FAQ. This is the first lever for gaining headroom.
- 03Close anything using VRAMA browser with hardware acceleration, a game, or video-editing software: each uses memory. On 6 GB, only one GPU workload at a time.
- 04Check with ollama psThe Processor column shows where the model was loaded: 100% GPU is the goal. Any split between CPU and GPU indicates overflow and lower speed.
- 05Choosing the quantizationQ4_K_M is the right compromise; on a 4B model, you can move up to Q5 or Q6 for better quality because you have room to spare. The quantization guide details the differences.
These variables are defined in the environment of the Ollama service (systemd on Linux, system variables on Windows, launchctl on macOS). Ollama's GPU error troubleshooting guide details what to check when a model refuses to load on the card.
#Diagnosis: why generation slows down on a 6 GB card
On a 6 GB card, a sudden slowdown almost always has the same cause: part of the model or cache has left the card. The following table links the observed symptom to the likely cause and the first response to try.
| Symptom | Likely cause | Try it |
|---|---|---|
| Smooth responses at first, very slow after a few exchanges | The context cache filled VRAM, and some layers spill over to the CPU | Reduce the context, quantize the cache to q8_0, restart the conversation |
| The model loads, but ollama ps shows a CPU/GPU split | The model and its context exceed the available memory | Choose a smaller model or a lighter quantization |
| Speed cut in half after opening a browser or a game | These applications use VRAM | Disable browser hardware acceleration, one GPU workload at a time |
| Memory error message while loading | The driver reserves memory and the model does not fit | Check the driver, free the card, try a 4B model |
| Disappointing throughput on a recent card with a narrow bus | Bandwidth is low (RTX 3050 6 GB: 96-bit bus) | Choose a smaller model instead of waiting for a software improvement |
One final point about quality: the smaller a model is, the more errors it makes on tasks that require precise knowledge or extended reasoning. With 6 GB, you’re better off using the model for what it does well (rewriting, summarizing provided text, answering from excerpts you give it, classifying, extracting fields) than asking it to answer factual questions from memory. Providing the source text in the conversation instead of relying on what the model “knows” is the practice that delivers the best results with a small model, and it remains compatible with a short context.
#Moving to 8 or 12 GB: what each tier unlocks
The most worthwhile jump is from 6 to 12 GB. At 12 GB, an 8-9B model in Q4 (5 to 6 GB) fits with a long context, and a 14B model in Q4 (about 9 GB) becomes possible. At 8 GB, an 8B model finally fits with room to spare, but a 14B model remains out of reach. Don't look for a price on this page: prices change every week, and the site's tracking reports the lowest price for each card.
| VRAM | Comfortable models | What changes |
|---|---|---|
| 6 GB | 3–4B, 8B with short context | A decent local assistant |
| 8 GB | 8B with medium context | More headroom, Q5 is possible |
| 12 GB | 8–9B with long context, 14B in Q4 | True comfort for chat and RAG |
| 16 GB | 14B in Q5, 24B in low Q3-Q4 | Higher-quality models |
- Local AI graphics card price tracking
- Which LLM for 8 GB of VRAM?
- Which LLM for 12 GB of VRAM
- Source: Ollama library, Qwen3
- Source: official Ollama FAQ
- Source: llama.cpp discussion about gpt-oss and --n-cpu-moe
#Frequently asked questions
What is the best LLM for 6 GB of VRAM?+
What speed should you expect on 6 GB of VRAM?+
Can you run Qwen 3.5 9B on 6 GB?+
RTX 2060 6 GB or RTX 3050 6 GB for local AI?+
Is 6 GB enough for local RAG?+
Can a 30B MoE run on 6 GB?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.