🇺🇸 Granite 4.1 8B Instruct
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Ranking updated on 09/10/2026
8 GB of VRAM is the entry point for local AI—RTX 3060 8GB, 4060, 5060, 3070, 2080, etc. 7–9B models in Q4_K_M fit comfortably. Here are the best choices for this VRAM budget.
RTX 4060 Ti 8GB : purchasing alternative available for local AI — RTX 5060 Ti 16 GB :
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Gemma 4 12B (Google): dense multimodal model (text, vision, audio), 256k context, ~7 GB Q4 VRAM. Apache 2.0, multilingual.
# HuggingFace : google/gemma-4-12B
Granite 4.2 8B (IBM): dense Apache 2.0, 128k context, ~4.6 GB Q4 VRAM. Multilingual chat, coding, and reasoning for the enterprise.
ollama pull granite4.2
SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.
ollama pull glm-5.3
Dense 7B 100% open (weights + data + code). Complete transparency for research.
ollama run olmo-3:7b
Hybrid thinking/fast mode. 119 languages, 32k native (131k via YaRN).
ollama run qwen3:8b
LFM2.5 7B (Liquid AI): dense Liquid Foundation Model, 32k context, 4.1 GB VRAM Q4. Optimized for CPU and edge. Released May 2026.
ollama pull lfm2.5
| Rank | Model | Params | Q4 VRAM | Context | License | On RTX 4060 Ti 8GB |
|---|---|---|---|---|---|---|
| #1 | Granite 4.1 8B Instruct | 8B | 5 GB | 131 072 | Apache 2.0 | 12 tok/s · Q5_K_M |
| #2 | Gemma 4 12B | 12B | 7 GB | 262 144 | Apache 2.0 | 18 tok/s · Q4_K_M |
| #3 | Granite 4.2 8B | 8B | 4.6 GB | 128 000 | Apache 2.0 | 32 tok/s · Q5_K_M |
| #4 | OLMo 3 7B Think (SFT) | 7B | 4.2 GB | 16 000 | Apache 2.0 | 32 tok/s · Q8 |
| #5 | GLM 5.3 7B | 7B | 4.1 GB | 128 000 | MIT | 32 tok/s · Q8 |
| #6 | OLMo 3 7B | 7B | 5 GB | 8 192 | Apache 2.0 | 12 tok/s · Q5_K_M |
| #7 | Qwen 3 8B | 8B | 5 GB | 131 072 | Apache 2.0 | 12 tok/s · Q5_K_M |
| #8 | LFM2.5 7B | 7B | 4.1 GB | 32 768 | LFM Open License v1.0 | 32 tok/s · Q8 |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: Q4 VRAM ≤ 8 GB and ≥ 3 GB (we exclude underused ultra-small models). We keep the 3–9B models that fit in Q4_K_M with room for context.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Which 8 GB LLM should you start with?
Granite 4.1 8B Instruct is the best starting point. Mistral 7B, Llama 3.1 8B, and Qwen 2.5 7B are all excellent and cost nothing. Start with Ollama (ollama run mistral:7b-instruct).
Can you run Gemma 2 9B in 8 GB?
Q4_K_M only (≈ 6 GB model + 1-2 GB context = ≈ 8 GB). No headroom for a large context. Prefer Mistral 7B if you want a comfortable 32k+ tokens.
Which quantization for 8 GB?
Q4_K_M for a 7–9B. Q5_K_M for a 3–5B. Q8 only for a 3B and below.
RTX 4060 vs RTX 3060 12 GB?
The 4060 is faster in raw throughput (+15-20%), but it's limited to 8 GB—no 12B models or large context. The 3060 12 GB is better for LLMs despite its lower tier.
Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.