Beginner 11 minBy VRAM

Which LLM for 6 GB of VRAM ?

Direct response

With 6 GB of VRAM, a 4-billion-parameter model in Q4 or Q5 fits comfortably with a context of several thousand tokens. A 7–8B model in Q4 (4.6 to 5.2 GB) still loads, but with no headroom: the context and memory reserved by the driver push it onto the CPU. Another surprising point: on 6 GB cards, memory bandwidth—not capacity—determines speed, and it varies from half to double.

6 GB is the threshold for GTX 1660, RTX 2060, and RTX 3050 6 GB. It's enough for a real local assistant, provided you choose the right model and account for memory that isn't used by the model. This page explains which models fit, which 6 GB card is faster than another, how to verify that a model stays on the card, and what 12 GB would change.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What 6 GB makes possible: the complete memory calculation

VRAM holds more than just the weights. It also contains the context cache, which grows with the conversation, the engine's computation buffers, and, under Windows, part of the display and browser workload. On a 6 GB card, you can budget approximately 5 to 5.5 GB for the model and its cache, with the rest taken by the system. The site's rule of thumb for Q4 weights is 3B ≈ 2 GB, 7-8B ≈ 5 GB.

What fits in 6 GB (Q4 size, excluding context)
ModelQ4 weightHeadroom above 5.5 GBVerdict
Gemma 4 2B1.2 GB4.3 GBVery large headroom, short and fast responses
Qwen 3.5 4B2.3 GB3.2 GBComfortable, with context of several thousand tokens
Phi-4 Mini 3.8B3 GB2.5 GBComfortable
Granite 4.2 8B4.6 GB0.9 GBSimply: very short context
Qwen 3 8B (Ollama)5.2 GB0.3 GBOverflows as soon as the context gets longer
Qwen 3.5 9B6 GBnegativeDoesn't fit

The sizes come from the QuelLLM catalog and the Ollama library. The “headroom” column is an estimate: it assumes 5.5 GB available, which depends on your system and applications. An 8B model is therefore possible on paper, but it leaves you with less than one gigabyte for the cache: the context setting determines whether the experience is smooth or the layers spill over onto the processor.

#Three 6 GB cards, three speeds: bandwidth

Generation speed is limited by memory bandwidth divided by the size of the weights read for each token. So two 6 GB cards are not equivalent. According to Wikipedia’s GeForce 20 generation table, the RTX 2060 6 GB delivers 336 GB/s on a 192-bit bus. The newer RTX 3050 6 GB has only a 96-bit bus (NVIDIA) and 168 GB/s according to the same encyclopedia: half as much.

Theoretical generation ceiling by card (bandwidth ÷ weights)
6 GB cardBandwidth2.3 GB model (4B Q4)4.6 GB model (8B Q4)
RTX 2060 6 GB336 GB/sapproximately 146 t/sabout 73 t/s
RTX 3050 6 GB168 GB/sabout 73 t/sabout 37 t/s

These ceilings are never reached, and we have no measurements to cite for these cards: the actual percentage depends on the engine, context, and card. But the 2-to-1 ratio holds. For local AI, a used RTX 2060 is therefore faster than a new 6 GB RTX 3050 despite its age, and the newer card has only energy efficiency and newer features in its favor. For GTX 1660, the variant matters: check the exact model number, because memory and bus width vary by model.

#The right model for every use case on 6 GB

Chat, questions, writing
A 3- to 4B model such as Qwen 3.5 4B or Phi-4 Mini: accurate, fast responses and comfortable context. For better French quality, try the 8B in Q4 with a reduced context.
Code
A small 3B to 7B coding model for autocompletion, and a general-purpose 4B model for explanations. Don’t expect a reliable coding agent at this size.
RAG and documents
A 4B model and a lightweight embedding model: everything fits in memory. Context is the limiting factor, so send short excerpts.
Vision and audio
Lightweight multimodal models exist, but the image encoder consumes memory in addition to the model: test with a 2 to 4B model.

A common trap: MoE models such as Qwen3-Coder 30B-A3B activate only 3 billion parameters, which may make you think they would fit on 6 GB. They do not: all experts must be in memory, and the file weighs 19 GB. The only way to use them is to keep the experts on the CPU, which requires a lot of RAM and comes at a steep speed cost. The guide to MoE models explains the principle.

#The hybrid GPU and CPU mode: useful, but not magic

When a model exceeds the card's capacity, Ollama and llama.cpp can place some layers on the CPU. This allows loading a 9B on 6 GB, but speed drops in stages: layers left in system RAM are read at system memory speed, a fraction of the card's speed. For MoE models, llama.cpp provides an option, --n-cpu-moe, which keeps the experts from certain layers on the CPU and leaves on the card what usually weighs the most. A user of the llama.cpp repository reports about 20 tokens per second with gpt-oss-20b and all experts on the CPU, on a machine with dual-channel DDR4-3200, while freeing most of the VRAM.

This account shows that the approach works, not that it suits your machine. It requires 32 GB of RAM for a model this size, and its throughput depends on memory bandwidth. For everyday use on 6 GB, a 4B model running entirely on the card remains simpler and faster.

#Five settings that matter on 6 GB

  1. 01
    Limit context
    Ollama's FAQ states that the default window is 4,096 tokens and can be adjusted. With 6 GB, increase it in increments (6,000, 8,000) and monitor memory: each increment adds cache.
  2. 02
    Quantize the K/V cache
    With Flash Attention enabled, OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache memory usage compared with f16, according to the Ollama FAQ. This is the first lever for gaining headroom.
  3. 03
    Close anything using VRAM
    A browser with hardware acceleration, a game, or video-editing software: each uses memory. On 6 GB, only one GPU workload at a time.
  4. 04
    Check with ollama ps
    The Processor column shows where the model was loaded: 100% GPU is the goal. Any split between CPU and GPU indicates overflow and lower speed.
  5. 05
    Choosing the quantization
    Q4_K_M is the right compromise; on a 4B model, you can move up to Q5 or Q6 for better quality because you have room to spare. The quantization guide details the differences.
Settings for a 6 GB card (set them, then relaunch Ollama)
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
OLLAMA_CONTEXT_LENGTH=6144

These variables are defined in the environment of the Ollama service (systemd on Linux, system variables on Windows, launchctl on macOS). Ollama's GPU error troubleshooting guide details what to check when a model refuses to load on the card.

#Diagnosis: why generation slows down on a 6 GB card

On a 6 GB card, a sudden slowdown almost always has the same cause: part of the model or cache has left the card. The following table links the observed symptom to the likely cause and the first response to try.

Common symptoms and remedies
SymptomLikely causeTry it
Smooth responses at first, very slow after a few exchangesThe context cache filled VRAM, and some layers spill over to the CPUReduce the context, quantize the cache to q8_0, restart the conversation
The model loads, but ollama ps shows a CPU/GPU splitThe model and its context exceed the available memoryChoose a smaller model or a lighter quantization
Speed cut in half after opening a browser or a gameThese applications use VRAMDisable browser hardware acceleration, one GPU workload at a time
Memory error message while loadingThe driver reserves memory and the model does not fitCheck the driver, free the card, try a 4B model
Disappointing throughput on a recent card with a narrow busBandwidth is low (RTX 3050 6 GB: 96-bit bus)Choose a smaller model instead of waiting for a software improvement

One final point about quality: the smaller a model is, the more errors it makes on tasks that require precise knowledge or extended reasoning. With 6 GB, you’re better off using the model for what it does well (rewriting, summarizing provided text, answering from excerpts you give it, classifying, extracting fields) than asking it to answer factual questions from memory. Providing the source text in the conversation instead of relying on what the model “knows” is the practice that delivers the best results with a small model, and it remains compatible with a short context.

#Moving to 8 or 12 GB: what each tier unlocks

The most worthwhile jump is from 6 to 12 GB. At 12 GB, an 8-9B model in Q4 (5 to 6 GB) fits with a long context, and a 14B model in Q4 (about 9 GB) becomes possible. At 8 GB, an 8B model finally fits with room to spare, but a 14B model remains out of reach. Don't look for a price on this page: prices change every week, and the site's tracking reports the lowest price for each card.

What each VRAM tier unlocks
VRAMComfortable modelsWhat changes
6 GB3–4B, 8B with short contextA decent local assistant
8 GB8B with medium contextMore headroom, Q5 is possible
12 GB8–9B with long context, 14B in Q4True comfort for chat and RAG
16 GB14B in Q5, 24B in low Q3-Q4Higher-quality models

#Frequently asked questions

FAQ
What is the best LLM for 6 GB of VRAM?+
A 3- to 4-billion-parameter model in Q4 or Q5, such as Qwen 3.5 4B (2.3 GB) or Phi-4 Mini: it fits with a comfortable context. An 8B model in Q4 (Granite 4.2 8B, 4.6 GB) is possible with a very short context, at the risk of spilling onto the CPU.
What speed should you expect on 6 GB of VRAM?+
It depends mainly on the card's memory bandwidth: 336 GB/s for a RTX 2060, 168 GB/s for a RTX 3050 6 GB. A 2.3 GB model has a theoretical ceiling of 146 t/s on the former and 73 t/s on the latter; real-world throughput is lower. We have no measurement to cite.
Can you run Qwen 3.5 9B on 6 GB?+
Not completely: it takes about 6 GB in Q4 in the QuelLLM catalog, which is the entire card with no room for the cache. In hybrid GPU and CPU mode, it loads but becomes slow. A 4B or 8B model with a short context is a better choice.
RTX 2060 6 GB or RTX 3050 6 GB for local AI?+
The RTX 2060: its 336 GB/s bandwidth is twice that of the RTX 3050 6 GB (168 GB/s), and it is what determines generation speed. The RTX 3050 6 GB wins only on power consumption and architectural generation, not throughput.
Is 6 GB enough for local RAG?+
Yes for simple RAG: a lightweight embedding model and a 4B model fit together in memory. The limiting factor is context length, which grows with the number of excerpts sent to the model. Limit them to a few short excerpts and quantize the K/V cache.
Can a 30B MoE run on 6 GB?+
Only by keeping the experts on the processor, which requires a great deal of RAM: the Qwen3-Coder 30B-A3B file weighs 19 GB. Despite its 3 billion active parameters, it does not fit on the card. This is an advanced approach, not the default configuration.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.