Which LLM for 8 GB of VRAM? ?
With 8 GB of VRAM, target an 8- to 9-billion-parameter model in Q4: Qwen 3.5 9B (6 GB of weights) or Granite 4.2 8B (4.6 GB), with a context of a few thousand tokens. A 12B model in Q4 barely fits; a 14B overflows. Keep about 1 GB free for the KV cache and the system.
Eight gigabytes of VRAM remains the most common capacity on cards sold over the past few years. This guide explains what fits, how to calculate your headroom, which cards actually have 8 GB, and when to move up to 12 or 16 GB.
Choosing a machine? Our picks by budget →
To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#With 8 GB of VRAM: 4- to 9-billion-parameter models
With 8 GB of VRAM, the right choice is an 8- to 9-billion-parameter model in Q4, such as Qwen 3.5 9B (6 GB of weights according to the QuelLLM catalog) or Granite 4.2 8B (4.6 GB), with a context of a few thousand tokens. A 12-billion-parameter model in Q4 (7 GB) barely fits, with no room left for context. A 14-billion-parameter model or larger exceeds the limit. What matters is not just the model weights: you must also add the KV cache, which grows with the context, and leave room for the system and display.
The budget calculation is simple: the model weights, plus the KV cache, plus about one gigabyte for buffers and display, must stay under 8 GB. With a 9B model in Q4 (6 GB), that leaves about 1 GB for context. With a lighter 8B model such as Granite 4.2 8B (4.6 GB), more than 2 GB remains. That's why the model and context choices must be made together.
| Model and quantization | Weights | Headroom for KV cache and system | Verdict |
|---|---|---|---|
| Granite 4.2 8B Q4 | 4.6 GB | ≈ 3.4 GB | Comfortable, long context possible |
| Granite 4.2 8B Q5 | 6 GB | ≈ 2 GB | Good compromise |
| Qwen 3.5 9B Q4 | 6 GB | ≈ 2 GB | The best general-purpose choice |
| Qwen 3.5 9B Q5 | 7 GB | ≈ 1 GB | Very short context |
| Gemma 4 12B Q4 | 7 GB | ≈ 1 GB | Just enough, with no margin |
| Qwen 3.5 4B Q4 | 2.3 GB | ≈ 5.7 GB | Very broad, for small tasks |
These margins are indicative: the actual KV cache size depends on the model architecture and context length, and you will get the best reading by launching the model and then checking ollama ps and your card’s monitoring tool (nvidia-smi or rocm-smi).
#Which cards have 8 GB, and which ones will trip you up
Many product lines come in multiple capacities, and the name is not enough. NVIDIA refers to the RTX 3060 with 12 GB or 8 GB, the RTX 4060 with 8 GB, the RTX 4060 Ti with 16 GB or 8 GB, and the RTX 5060 Ti with 16 GB or 8 GB, while the RTX 5060 comes with 8 GB. On the AMD side, the RX 9060 XT is available with 8 GB, with 320 GB/s of bandwidth like the 16 GB version. Always verify the capacity on the exact model’s specifications page before buying or comparing.
- NVIDIA 8 GB
- RTX 4060, RTX 5060, 8 GB versions of the RTX 3060, 4060 Ti and 5060 Ti, and the older RTX 3070, 3060 Ti, or 2060 Super.
- AMD 8 GB
- RX 9060 XT 8 GB, RX 7600, RX 6600 XT, RX 5700 XT.
- Warning: RX 7600 XT
- This card has 16 GB, not 8: the llama.cpp performance table lists it with 16 GB of GDDR6. An earlier version of this guide incorrectly classified it as 8 GB.
On Mac, memory is unified and shared with the system: an 8 GB MacBook Air does not reserve 8 GB for AI. The guides dedicated to the MacBook Air M1 and M2 explain what remains available; with no dedicated graphics card at all, the guide to LLMs without a GPU lists models by RAM capacity.
#What speed to expect on 8 GB
Text generation reads nearly all the weights at each token: maximum speed is memory bandwidth divided by weight size. On an RX 9060 XT 8 GB (320 GB/s according to AMD), the theoretical ceiling is about 53 tokens per second for an Qwen 3.5 9B in Q4 (6 GB) and 70 for an Granite 4.2 8B in Q4 (4.6 GB). Another card will have a different ceiling; actual figures are always lower.
We do not publish tokens per second by card: they depend on the driver, software, and context, and no public benchmark we can cite covers these models on these cards. The llama.cpp community table, which measures a Llama 2 7B in Q4_0, provides a reference point for an older 8 GB card: 67 tokens/s generated on a RX 5700 XT under ROCm. For your hardware, measure it: ollama run with the --verbose option displays the generation speed.
#Which models to choose on 8 GB depending on the use case
The QuelLLM catalog helps you identify a model with headroom for each use case. The largest context advertised by a model does not mean you will be able to use it: on 8 GB, the free memory left after the weights—not the catalog figure—limits the actual length.
| Usage | Model | Q4 weight | Note |
|---|---|---|---|
| Discussion and writing | Qwen 3.5 9B | 6 GB | Advertised context of 262,000 tokens; practically usable well below that on 8 GB |
| Everyday tasks, plenty of headroom | Granite 4.2 8B | 4.6 GB | Advertised context of 128,000 tokens; leaves more room for the KV cache |
| Code completion | Qwen 2.5 Coder 7B | 5 GB | 7-billion-parameter code model, stated context of 131,072 tokens |
| Small background tasks | Qwen 3.5 4B | 2.3 GB | Leaves VRAM free for other tools |
| Vision and multimodal | Gemma 4 12B | 7 GB | Tagged for vision and audio; barely fits in 8 GB |
Two useful details. Models labeled vision use additional memory for the image encoder, so a 12-billion-parameter multimodal model in Q4 quickly exceeds 8 GB as soon as an image is loaded. For code, a specialized 7-billion-parameter model in Q4 (5 GB) fits with room to spare and responds faster than a larger general-purpose model, which matters for online completion.
#Five settings to stretch 8 GB
None of these settings replaces VRAM, but they push the limit back by a few hundred megabytes to one gigabyte, which is enough to go from a 4 000-token context to a 16 000-token context.
- Limit context
- According to its documentation, Ollama uses a 4,096-token context window by default; OLLAMA_CONTEXT_LENGTH changes it. Don't open 32,000 tokens if you're chatting in short messages.
- Quantize the KV cache
- OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache memory usage compared with the default f16; it requires Flash Attention.
- Let Flash Attention do its job
- Ollama uses it automatically when the backend and card support it; you can force it with OLLAMA_FLASH_ATTENTION=1.
- Free the card
- Browsers, games, and video applications consume VRAM. Close them before loading a model, or connect the display to the integrated graphics processor if one is available.
- Choose Q4 or Q5 depending on your headroom
- On 8 GB, Q4_K_M is the right default. Move to Q5 only with a smaller model that leaves more than one gigabyte free.
#What 8 GB doesn't allow: RAG, code, large models
A complete RAG pipeline combines an embedding model, optionally a reranker, and the LLM that writes the answer. On 8 GB, this setup does not fit comfortably: the generation model alone takes 5 to 6 GB. Two solutions: run the embedder on the CPU, which is entirely acceptable for indexing, or choose a small generation model such as Qwen 3.5 4B in Q4 (2.3 GB). For online code completion, a 7-billion-parameter model is enough; for a coding agent that reads large files, the required context exceeds what 8 GB allows.
For models with 14 to 32 billion parameters, you need to move up a tier. The site’s rule of thumb is 14B ≈ 9 GB and 32B ≈ 19 to 20 GB, in Q4 and excluding context: 14B is therefore the real boundary above 8 GB, which is why a 12 GB tier is useful—it accommodates it with headroom. Do not count on extreme offloading tricks: a model that fits half in RAM may run, but with read speeds that make interactive use discouraging.
#When to move to 12 or 16 GB
| Tier | What becomes possible | Guide |
|---|---|---|
| 12 GB | Gemma 4 12B in Q8, Qwen 3.5 9B in Q8, lighter RAG | Which LLM for 12 GB of VRAM |
| 16 GB | 20B to 24B models in Q4 (gpt-oss 20B, Mistral Small 24B) with some context | Which LLM for 16 GB of VRAM |
| 24 GB | 27-billion-parameter dense models in Q4 with a long context | Which LLM for 24 GB of VRAM |
The right purchasing criterion is the price per gigabyte of VRAM, not the card’s price. The site’s price tracker calculates it for every card and updates twice a week; check it before deciding, because the gap between new and used cards often changes.
- Which LLM for 12 GB of VRAM
- Which LLM for 16 GB of VRAM
- Quantize the KV cache to save VRAM
- Choose your quantization (Q4, Q5, Q8, FP16)
- Local LLM without a GPU: models by RAM
- Local AI graphics card prices
- Source: Ollama FAQ (context, KV cache, Flash Attention)
- Source: NVIDIA model card from the RTX 5060 family
- Source: AMD spec sheet for the RX 9060 XT 8 GB
- QuelLLM model catalog (VRAM sizes)
#Frequently asked questions
What’s the best LLM for 8 GB of VRAM?+
Can you run a 14B model on 8 GB?+
How can I tell whether the model fits entirely in my VRAM?+
Does the RX 7600 XT have 8 GB?+
Is 8 GB enough to develop an application with an LLM?+
Is a Mac with 8 GB equivalent to a PC with 8 GB of VRAM?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.