Home › Catalog › Best LLM for 8 GB of VRAM in 2026

Best LLM for 8 GB of VRAM in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

8 GB of VRAM is the entry point for local AI—RTX 3060 8GB, 4060, 5060, 3070, 2080, etc. 7–9B models in Q4_K_M fit comfortably. Here are the best choices for this VRAM budget.

Offers and alternatives for local AI

RTX 4060 Ti 8GB : purchasing alternative available for local AI — RTX 5060 Ti 16 GB :

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇺🇸 Granite 4.1 8B Instruct

IBM · 8B parameters · Apache 2.0 · 131,072 tokens ctx

Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.

Why this ranking Fits in Q5_K_M (~6 GB out of 8 GB available). 8B parameters, 131,072-token context.
# HuggingFace : ibm-granite/granite-4.1-8b
On RTX 4060 Ti 8GB
Q5_K_M
6 GB · 12 tok/s
2

🇺🇸 Gemma 4 12B

Google · 12B parameters · Apache 2.0 · 262,144-token context

Gemma 4 12B (Google): dense multimodal model (text, vision, audio), 256k context, ~7 GB Q4 VRAM. Apache 2.0, multilingual.

Why this ranking Fits in Q4_K_M (~7 GB out of 8 GB available). 12B parameters, 262,144-token context.
# HuggingFace : google/gemma-4-12B
On RTX 4060 Ti 8GB
Q4_K_M
7 GB · 18 tok/s
3

🇺🇸 Granite 4.2 8B

IBM · 8B parameters · Apache 2.0 · 128,000 tokens ctx

Granite 4.2 8B (IBM): dense Apache 2.0, 128k context, ~4.6 GB Q4 VRAM. Multilingual chat, coding, and reasoning for the enterprise.

Why this ranking Fits in Q5_K_M (~6 GB out of 8 GB available). 8B parameters, 128,000-token context.
ollama pull granite4.2
On RTX 4060 Ti 8GB
Q5_K_M
6 GB · 32 tok/s
4

🇺🇸 OLMo 3 7B Think (SFT)

zimplex · 7B parameters · Apache 2.0 · 16,000 tokens ctx

SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.

Why this ranking Fits in Q8 (~8 GB out of 8 GB available). 7B parameters, 16,000-token context.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
On RTX 4060 Ti 8GB
Q8
8 GB · 32 tok/s
5

🇨🇳 GLM 5.3 7B

Zhipu AI · 7B parameters · MIT · 128,000 tokens ctx

GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.

Why this ranking Fits in Q8 (~7 GB out of 8 GB available). 7B parameters, 128,000-token context.
ollama pull glm-5.3
On RTX 4060 Ti 8GB
Q8
7 GB · 32 tok/s
6

🇺🇸 OLMo 3 7B

Allen AI · 7B parameters · Apache 2.0 · 8,192-token context

Dense 7B 100% open (weights + data + code). Complete transparency for research.

Why this ranking Fits in Q5_K_M (~6 GB of 8 GB available). 7B parameters, 8,192-token context.
ollama run olmo-3:7b
On RTX 4060 Ti 8GB
Q5_K_M
6 GB · 12 tok/s
7

🇨🇳 Qwen 3 8B

Alibaba · 8B parameters · Apache 2.0 · 131,072 tokens ctx

Hybrid thinking/fast mode. 119 languages, 32k native (131k via YaRN).

Why this ranking Fits in Q5_K_M (~6 GB out of 8 GB available). 8B parameters, 131,072-token context.
ollama run qwen3:8b
On RTX 4060 Ti 8GB
Q5_K_M
6 GB · 12 tok/s
8

🇺🇸 LFM2.5 7B

Liquid AI · 7B parameters · LFM Open License v1.0 · 32,768-token context

LFM2.5 7B (Liquid AI): dense Liquid Foundation Model, 32k context, 4.1 GB VRAM Q4. Optimized for CPU and edge. Released May 2026.

Why this ranking Fits in Q8 (~7 GB out of 8 GB available). 7B parameters, 32 768-token context.
ollama pull lfm2.5
On RTX 4060 Ti 8GB
Q8
7 GB · 32 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On RTX 4060 Ti 8GB
#1 Granite 4.1 8B Instruct 8B 5 GB 131 072 Apache 2.0 12 tok/s · Q5_K_M
#2 Gemma 4 12B 12B 7 GB 262 144 Apache 2.0 18 tok/s · Q4_K_M
#3 Granite 4.2 8B 8B 4.6 GB 128 000 Apache 2.0 32 tok/s · Q5_K_M
#4 OLMo 3 7B Think (SFT) 7B 4.2 GB 16 000 Apache 2.0 32 tok/s · Q8
#5 GLM 5.3 7B 7B 4.1 GB 128 000 MIT 32 tok/s · Q8
#6 OLMo 3 7B 7B 5 GB 8 192 Apache 2.0 12 tok/s · Q5_K_M
#7 Qwen 3 8B 8B 5 GB 131 072 Apache 2.0 12 tok/s · Q5_K_M
#8 LFM2.5 7B 7B 4.1 GB 32 768 LFM Open License v1.0 32 tok/s · Q8
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: Q4 VRAM ≤ 8 GB and ≥ 3 GB (we exclude underused ultra-small models). We keep the 3–9B models that fit in Q4_K_M with room for context.

Criteria considered:

  • Q4 VRAM ≤ 8 GB
  • Usable context (32k+)
  • Comparable quality to the top tier
  • Compatible with Ollama / LM Studio

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

Which 8 GB LLM should you start with?

Granite 4.1 8B Instruct is the best starting point. Mistral 7B, Llama 3.1 8B, and Qwen 2.5 7B are all excellent and cost nothing. Start with Ollama (ollama run mistral:7b-instruct).

Can you run Gemma 2 9B in 8 GB?

Q4_K_M only (≈ 6 GB model + 1-2 GB context = ≈ 8 GB). No headroom for a large context. Prefer Mistral 7B if you want a comfortable 32k+ tokens.

Which quantization for 8 GB?

Q4_K_M for a 7–9B. Q5_K_M for a 3–5B. Q8 only for a 3B and below.

RTX 4060 vs RTX 3060 12 GB?

The 4060 is faster in raw throughput (+15-20%), but it's limited to 8 GB—no 12B models or large context. The 3060 12 GB is better for LLMs despite its lower tier.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.