Beginner 8 minBy VRAM

Which LLM for 4 GB of VRAM ?

4 GB of VRAM is the absolute minimum for running an LLM locally in 2026. GTX 1650, GTX 1050 Ti, RTX 3050 Mobile… older or ultra-entry-level cards. You can still install Ollama and chat with Qwen 3.5 2B or Granite 4.2 3B. This guide explains exactly what works, what does not, and how to get the most out of those 4 GB.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

To move to 16 GB of VRAM: RTX 5060 Ti 16GB (ASUS Prime).

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#4 GB of VRAM in 2026

i
The bare minimum
4 GB of VRAM means you can test an LLM locally, but with serious limitations. An 8B model from 2026 such as Granite 4.2 8B Q4 (5.3 GB) exceeds that—stick with comfortable 2B–4B models.

#1. Which GPUs Have 4 GB

GTX 1650 (4 GB)
Pascal-Turing 2019. Difficult to find new today; used, prices vary.
GTX 1050 Ti (4 GB)
Pascal 2016. Old but still functional. Used, with variable pricing.
RTX 3050 Mobile 4 GB
Entry-level laptop variant.
GTX 1630 (4 GB)
Ultra-entry-level, avoid—it performs poorly even for lightweight LLMs.

#2. Usable models

4 GB of VRAM — 2026 LLM models
ModelQuantVRAMOK?
Gemma 4 2BQ5_K_M1.4 GB★★★★★ comfortable
Qwen 3.5 2BQ4_K_M1.9 GB★★★★★ comfortable
Granite 4.2 3BQ4_K_M2.2 GB★★★★★
Qwen 3.5 4BQ4_K_M2.3 GB★★★★
Gemma 4 E2BQAT4.3 GB⚠️ tangent (4 GB limit)
Granite 4.2 8BQ4_K_M5.3 GB❌ overflows

#3. Tips for stretching 4 GB

Q4 KV cache
OLLAMA_KV_CACHE_TYPE=q4_0 (instead of q8_0). More aggressive, it saves 50% of the cache.
2k context max
Beyond that, everything spills over. 2k tokens = ~5–10 sentences in French, enough for simple chat.
Close everything else
GPU-accelerated Chrome, Discord, OBS—all consume VRAM. Minimalist mode is mandatory.
No RAG
Embeddings + LLM = too much. Stick to simple chat without documents.

#Recommended upgrade

RTX 3060 12 GB (used, variable price)
The most cost-effective leap. +200% VRAM, access to 14B models. The best possible LLM upgrade.
RTX 5050 8 GB (new price to verify, not tracked)
If you want something new. +100% VRAM, future-proof FP4.
Mac mini M4 16 GB (generally replaced by M5/M6; verify)
If you switch platforms. Unified memory, quiet.

#Frequently asked questions

Can you run a 7-8B model with 4 GB of VRAM?+
Not comfortably. A 2026 8B model like Granite 4.2 8B Q4 (5.3 GB) overflows. Stick with Granite 4.2 3B Q4 (2.2 GB) — decent quality, decent speed.
What’s the best LLM for 4 GB of VRAM?+
Granite 4.2 3B Q4_K_M or Qwen 3.5 4B Q4. Excellent small models for 2026, with much better quality than a 7B model from 2 years ago.
Is GTX 1650 4 GB enough to explore local AI?+
Yes for 2B–3B models. Install Ollama, then chat with Granite 4.2 3B at 25–30 tok/s. Enough to get started.
Do you absolutely need a GPU? Can my CPU be enough?+
On a recent CPU (Ryzen 5 7600+, Core i5 13th gen+), Granite 4.2 8B Q4 runs at 5–10 tok/s. Slow but usable. With 4 GB of GPU memory, you get ~20 tok/s on 3B—preferable.
Can you do RAG with 4 GB of VRAM?+
Not comfortably. BGE-M3 embeddings (~2 GB) + 3B LLM (~2 GB) = saturated. Stick to simple chat or move to 8 GB+.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.