🇺🇸 Granite 4.1 8B Instruct
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Ranking updated on 09/10/2026
The RTX 3080 10 GB (GDDR6X, 760 GB/s) remains highly capable in 2026. 10 GB limits 13B models in Q4 (~8 GB leaves little room for context), but 7–9B Q5 and Mistral 7B FP16 run excellently.
RTX 3080 10GB : purchasing alternative available for local AI — RTX 5070 Ti 16 GB :
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Granite 4.2 8B (IBM): dense Apache 2.0, 128k context, ~4.6 GB Q4 VRAM. Multilingual chat, coding, and reasoning for the enterprise.
ollama pull granite4.2
SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.
ollama pull glm-5.3
Next-generation dense 9B. 262k ctx, improved hybrid thinking.
ollama run qwen3.5:9b
Dense 8B vision Qwen 3. Best small VLM Qwen generation 3.
ollama run qwen3-vl:8b
Compact 70B version. 1000+ languages, trained on Swiss Alps supercomputer.
ollama pull hf.co/swissai/Apertus-8B-GGUF
| Rank | Model | Params | Q4 VRAM | Context | License | On RTX 3080 10GB |
|---|---|---|---|---|---|---|
| #1 | Granite 4.1 8B Instruct | 8B | 5 GB | 131 072 | Apache 2.0 | 35 tok/s · Q8 |
| #2 | Granite 4.2 8B | 8B | 4.6 GB | 128 000 | Apache 2.0 | 50 tok/s · Q8 |
| #3 | OLMo 3 7B Think (SFT) | 7B | 4.2 GB | 16 000 | Apache 2.0 | 50 tok/s · Q8 |
| #4 | GLM 5.3 7B | 7B | 4.1 GB | 128 000 | MIT | 50 tok/s · Q8 |
| #5 | Qwen 3.5 9B | 9B | 6 GB | 262 000 | Apache 2.0 | 28 tok/s · Q8 |
| #6 | Qwen 3 VL 8B | 8B | 6 GB | 262 144 | Apache 2.0 | 30 tok/s · Q8 |
| #7 | Apertus 8B | 8B | 6 GB | 65 536 | Apache 2.0 | 30 tok/s · Q8 |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: Q4_K_M ≤ 9 GB. Bonus: 7–9B (10 GB peak). 760 GB/s = solid.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
RTX 3080 10 GB in 2026: still relevant?
Yes for 7-9B in Q5/Q8 (Mistral 7B Q8 = ~7.5 GB) at 50+ tok/s. For 13-14B, you need Q4 with little headroom. See guide.
3080 10 GB vs. 4070 12 GB?
3080 = 760 GB/s, 4070 = 504 GB/s. 3080 ~50% faster. But the 4070 has 12 GB (Qwen 3 14B Q4 OK). Depending on whether speed or VRAM is the priority. See RTX 4070.
Which quantization for a 3080 10 GB?
Q8 for 7B (near-FP16 quality, ~7.5 GB). Q5_K_M for 8–9B (~6–7 GB). Avoid Q4 unless you want to attempt a 13B (~8 GB, little headroom).
Used 3080 10 GB: how much?
~€350-400 in France. For LLMs, a used 3090 (~€650) remains better if you're on a budget. See RTX 3090.
Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.