Beginner 11 minBy VRAM

Which LLM for 8 GB of VRAM? ?

Direct response

With 8 GB of VRAM, target an 8- to 9-billion-parameter model in Q4: Qwen 3.5 9B (6 GB of weights) or Granite 4.2 8B (4.6 GB), with a context of a few thousand tokens. A 12B model in Q4 barely fits; a 14B overflows. Keep about 1 GB free for the KV cache and the system.

Eight gigabytes of VRAM remains the most common capacity on cards sold over the past few years. This guide explains what fits, how to calculate your headroom, which cards actually have 8 GB, and when to move up to 12 or 16 GB.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#With 8 GB of VRAM: 4- to 9-billion-parameter models

With 8 GB of VRAM, the right choice is an 8- to 9-billion-parameter model in Q4, such as Qwen 3.5 9B (6 GB of weights according to the QuelLLM catalog) or Granite 4.2 8B (4.6 GB), with a context of a few thousand tokens. A 12-billion-parameter model in Q4 (7 GB) barely fits, with no room left for context. A 14-billion-parameter model or larger exceeds the limit. What matters is not just the model weights: you must also add the KV cache, which grows with the context, and leave room for the system and display.

The budget calculation is simple: the model weights, plus the KV cache, plus about one gigabyte for buffers and display, must stay under 8 GB. With a 9B model in Q4 (6 GB), that leaves about 1 GB for context. With a lighter 8B model such as Granite 4.2 8B (4.6 GB), more than 2 GB remains. That's why the model and context choices must be made together.

Memory budget on 8 GB (weights: QuelLLM catalog)
Model and quantizationWeightsHeadroom for KV cache and systemVerdict
Granite 4.2 8B Q44.6 GB≈ 3.4 GBComfortable, long context possible
Granite 4.2 8B Q56 GB≈ 2 GBGood compromise
Qwen 3.5 9B Q46 GB≈ 2 GBThe best general-purpose choice
Qwen 3.5 9B Q57 GB≈ 1 GBVery short context
Gemma 4 12B Q47 GB≈ 1 GBJust enough, with no margin
Qwen 3.5 4B Q42.3 GB≈ 5.7 GBVery broad, for small tasks

These margins are indicative: the actual KV cache size depends on the model architecture and context length, and you will get the best reading by launching the model and then checking ollama ps and your card’s monitoring tool (nvidia-smi or rocm-smi).

i
A correction to the size estimates
An earlier version of this guide listed 6.6 GB for Qwen 3.5 9B and 7.6 GB for Gemma 4 12B in Q4. The QuelLLM catalog gives 6 and 7 GB. These values are rounded: always keep one gigabyte of headroom over what you calculate.

#Which cards have 8 GB, and which ones will trip you up

Many product lines come in multiple capacities, and the name is not enough. NVIDIA refers to the RTX 3060 with 12 GB or 8 GB, the RTX 4060 with 8 GB, the RTX 4060 Ti with 16 GB or 8 GB, and the RTX 5060 Ti with 16 GB or 8 GB, while the RTX 5060 comes with 8 GB. On the AMD side, the RX 9060 XT is available with 8 GB, with 320 GB/s of bandwidth like the 16 GB version. Always verify the capacity on the exact model’s specifications page before buying or comparing.

NVIDIA 8 GB
RTX 4060, RTX 5060, 8 GB versions of the RTX 3060, 4060 Ti and 5060 Ti, and the older RTX 3070, 3060 Ti, or 2060 Super.
AMD 8 GB
RX 9060 XT 8 GB, RX 7600, RX 6600 XT, RX 5700 XT.
Warning: RX 7600 XT
This card has 16 GB, not 8: the llama.cpp performance table lists it with 16 GB of GDDR6. An earlier version of this guide incorrectly classified it as 8 GB.

On Mac, memory is unified and shared with the system: an 8 GB MacBook Air does not reserve 8 GB for AI. The guides dedicated to the MacBook Air M1 and M2 explain what remains available; with no dedicated graphics card at all, the guide to LLMs without a GPU lists models by RAM capacity.

#What speed to expect on 8 GB

Text generation reads nearly all the weights at each token: maximum speed is memory bandwidth divided by weight size. On an RX 9060 XT 8 GB (320 GB/s according to AMD), the theoretical ceiling is about 53 tokens per second for an Qwen 3.5 9B in Q4 (6 GB) and 70 for an Granite 4.2 8B in Q4 (4.6 GB). Another card will have a different ceiling; actual figures are always lower.

We do not publish tokens per second by card: they depend on the driver, software, and context, and no public benchmark we can cite covers these models on these cards. The llama.cpp community table, which measures a Llama 2 7B in Q4_0, provides a reference point for an older 8 GB card: 67 tokens/s generated on a RX 5700 XT under ROCm. For your hardware, measure it: ollama run with the --verbose option displays the generation speed.

Measure your own speed
ollama run <nom-du-modèle> --verbose
!
When the model spills over, speed drops
If part of the model does not fit in VRAM, Ollama places it in RAM and on the processor. Generation then runs at RAM speed, often several times slower. ollama ps shows the split: aim for 100% GPU.

#Which models to choose on 8 GB depending on the use case

The QuelLLM catalog helps you identify a model with headroom for each use case. The largest context advertised by a model does not mean you will be able to use it: on 8 GB, the free memory left after the weights—not the catalog figure—limits the actual length.

A starting point by use case (QuelLLM catalog)
UsageModelQ4 weightNote
Discussion and writingQwen 3.5 9B6 GBAdvertised context of 262,000 tokens; practically usable well below that on 8 GB
Everyday tasks, plenty of headroomGranite 4.2 8B4.6 GBAdvertised context of 128,000 tokens; leaves more room for the KV cache
Code completionQwen 2.5 Coder 7B5 GB7-billion-parameter code model, stated context of 131,072 tokens
Small background tasksQwen 3.5 4B2.3 GBLeaves VRAM free for other tools
Vision and multimodalGemma 4 12B7 GBTagged for vision and audio; barely fits in 8 GB

Two useful details. Models labeled vision use additional memory for the image encoder, so a 12-billion-parameter multimodal model in Q4 quickly exceeds 8 GB as soon as an image is loaded. For code, a specialized 7-billion-parameter model in Q4 (5 GB) fits with room to spare and responds faster than a larger general-purpose model, which matters for online completion.

#Five settings to stretch 8 GB

None of these settings replaces VRAM, but they push the limit back by a few hundred megabytes to one gigabyte, which is enough to go from a 4 000-token context to a 16 000-token context.

Limit context
According to its documentation, Ollama uses a 4,096-token context window by default; OLLAMA_CONTEXT_LENGTH changes it. Don't open 32,000 tokens if you're chatting in short messages.
Quantize the KV cache
OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache memory usage compared with the default f16; it requires Flash Attention.
Let Flash Attention do its job
Ollama uses it automatically when the backend and card support it; you can force it with OLLAMA_FLASH_ATTENTION=1.
Free the card
Browsers, games, and video applications consume VRAM. Close them before loading a model, or connect the display to the integrated graphics processor if one is available.
Choose Q4 or Q5 depending on your headroom
On 8 GB, Q4_K_M is the right default. Move to Q5 only with a smaller model that leaves more than one gigabyte free.
Ollama with quantized KV cache
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=8192 ollama serve

#What 8 GB doesn't allow: RAG, code, large models

A complete RAG pipeline combines an embedding model, optionally a reranker, and the LLM that writes the answer. On 8 GB, this setup does not fit comfortably: the generation model alone takes 5 to 6 GB. Two solutions: run the embedder on the CPU, which is entirely acceptable for indexing, or choose a small generation model such as Qwen 3.5 4B in Q4 (2.3 GB). For online code completion, a 7-billion-parameter model is enough; for a coding agent that reads large files, the required context exceeds what 8 GB allows.

For models with 14 to 32 billion parameters, you need to move up a tier. The site’s rule of thumb is 14B ≈ 9 GB and 32B ≈ 19 to 20 GB, in Q4 and excluding context: 14B is therefore the real boundary above 8 GB, which is why a 12 GB tier is useful—it accommodates it with headroom. Do not count on extreme offloading tricks: a model that fits half in RAM may run, but with read speeds that make interactive use discouraging.

#When to move to 12 or 16 GB

What do you gain by moving up in capacity
TierWhat becomes possibleGuide
12 GBGemma 4 12B in Q8, Qwen 3.5 9B in Q8, lighter RAGWhich LLM for 12 GB of VRAM
16 GB20B to 24B models in Q4 (gpt-oss 20B, Mistral Small 24B) with some contextWhich LLM for 16 GB of VRAM
24 GB27-billion-parameter dense models in Q4 with a long contextWhich LLM for 24 GB of VRAM

The right purchasing criterion is the price per gigabyte of VRAM, not the card’s price. The site’s price tracker calculates it for every card and updates twice a week; check it before deciding, because the gap between new and used cards often changes.

#Frequently asked questions

Frequently asked questions
What’s the best LLM for 8 GB of VRAM?+
Qwen 3.5 9B in Q4, with about 6 GB of weights according to the QuelLLM catalog, is the best general-purpose model, with a context of a few thousand tokens. Granite 4.2 8B in Q4 (4.6 GB) leaves more room for a long context. For code completion, a specialized 7-billion-parameter model is sufficient.
Can you run a 14B model on 8 GB?+
Not comfortably. The site's rule gives about 9 GB of weights for a 14B model in Q4, which is more than the card has. Some of it would spill onto the processor, and speed would collapse. Gemma 4 12B in Q4 (7 GB) barely fits, with no headroom for the context: choose a 9B model instead.
How can I tell whether the model fits entirely in my VRAM?+
Generate a response, then run ollama ps: the PROCESSOR column shows the split between GPU and processor, and your target is 100% GPU. If some work shifts to the processor, reduce the context, quantize the KV cache with OLLAMA_KV_CACHE_TYPE, or choose a smaller model.
Does the RX 7600 XT have 8 GB?+
No, it has 16 GB of GDDR6, as shown in the llama.cpp performance table. This is a common mistake because the base RX 7600 has 8 GB. Always check the model's exact capacity: several product lines come in 8 GB and 16 GB versions, such as the RTX 5060 Ti and the RX 9060 XT.
Is 8 GB enough to develop an application with an LLM?+
For prototyping a chat or extraction app with an 8B model, yes. For a production RAG system with an embedder, reranker, and a 14-billion-parameter model, no: target 12 to 16 GB. You can also run the embedder on the CPU to free VRAM for the LLM.
Is a Mac with 8 GB equivalent to a PC with 8 GB of VRAM?+
No. On a Mac, memory is unified and shared with the system and applications, so the amount actually available to the model is less than 8 GB. On a PC with a dedicated card, the 8 GB of VRAM is reserved for the GPU, leaving more room for the model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.