Gemma 4 locally with Ollama: installation, VRAM, perfs
Gemma 4 is Google's new generation of open-weight models, released in 2026. Native multimodality, long context, an Apache 2.0 license, and three useful sizes in the Ollama library: E2B (compact), 12B, and 26B-A4B (a MoE). This guide shows you how to run Gemma 4 with Ollama, which quantization to choose based on your VRAM, approximate tokens/sec on RTX and Apple Silicon, and what concretely changes compared with Gemma 3.
#Why run Gemma 4 locally
Gemma 4 is one of the best open-weight models of its size for French. The 12B Q4 fits comfortably in 8 to 12 GB of VRAM, covering a RTX 3060, a 4070, or a midrange M-series Mac. The 26B-A4B, an MoE that activates only 4B parameters, makes use of 24 GB cards (3090, 4090, 7900 XTX) and Macs with plenty of unified memory while remaining fast. For a French-language assistant, RAG, and general-purpose coding, it's an excellent default.
The other reason to run it locally: since April 2026 Gemma 4 has been released under the Apache 2.0 license, lifting the restrictions of the former Gemma Terms of Use (free commercial use, redistribution, and fine-tuning), and Ollama handles pulling, quantization, and inference without requiring you to touch Python, CUDA, or llama.cpp directly.
#What changes versus Gemma 3
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Gemma 3 (March 2025) introduced multimodality and a 128k context window. Gemma 4 preserves those gains and pushes forward on three fronts: quality (a major measured leap on reasoning and coding benchmarks), efficiency (the 12B replaces Gemma 3's 27B to good effect on many tasks, and the 26B-A4B MoE runs the equivalent of a large model at the speed of a small one), and tokenizer (a denser vocabulary, which reduces the number of tokens for equivalent output—so more useful tokens per second).
- Sizes
- Gemma 3: 1B, 4B, 12B, 27B (dense). Gemma 4: E2B (compact), 12B (dense), and 26B-A4B (MoE, 4B active). The lineup scales in efficiency rather than raw size.
- Multimodal
- Native vision on the 12B and 26B-A4B (text + images). The E2B remains the compact, text-focused variant.
- Context
- 128k tokens across the entire lineup. Same memory constraints as before: the KV cache adds up quickly.
- Tokenizer
- SentencePiece, with a vocabulary revised to better cover French, code, and non-Latin languages. With equivalent prompts, it uses ~5 to 15% fewer tokens.
- License
- Apache 2.0 since April 2026. Commercial use, redistribution, and fine-tuning are permitted, with no specific acceptable-use clause—a real change from the Gemma Terms of Use of Gemma 3.
#Prerequisites
- Ollama installed
- Recent version (≥ 0.5). See the Windows, macOS, or Linux installation guides if you haven't already.
- Disk space
- Allow 4.3 GB for the E2B, 7.6 GB for the 12B Q4, and 19 GB for the 26B-A4B Q4. The Q5 and Q8 variants of the 12B and 26B weigh 1.3 to 2× more.
- VRAM or unified memory
- Practical minimum: 6 GB for the E2B, 8 to 12 GB for the 12B in Q4, 24 GB for the 26B-A4B in Q4. Below that, it spills onto the CPU and speed collapses.
- Up-to-date GPU drivers
- CUDA 12.x for NVIDIA, ROCm 6.x for AMD, native Metal for Apple Silicon. Ollama detects the backend automatically.
#1. Pull and launch with one command
The Ollama tag for Gemma 4 follows the usual convention: gemma4 (latest points to the 12B Q4 by default), gemma4:e2b-it-qat, gemma4:12b, gemma4:26b. To specify quantization, add a suffix: gemma4:12b-instruct-q4_K_M.
On the first run, Ollama downloads the model (~7.6 GB for the 12B Q4), then opens an interactive prompt directly. Subsequent launches are instantaneous as long as the model remains in memory cache.
To check what is currently loaded in VRAM:
The PROCESSOR column must show 100% GPU. If you see 70%/30% GPU/CPU, the model is spilling over—switch to more aggressive quantization or a smaller size.
#2. VRAM by size and quantization
The figures below include the model weights plus a KV cache for an 8k-token context window (the Ollama default). To push to 32k or 128k, budget an additional 2 to 8 GB depending on size.
- Gemma 4 E2B — QAT int4
- ≈ 4.5 GB VRAM. Compact variant (QAT, ~2B active parameters in the MatFormer style). Runs on any GTX 1660 (6 GB), RTX 3050 (8 GB), or Mac with 8 GB of unified memory.
- Gemma 4 12B — Q4_K_M
- ≈ 8 GB of VRAM. Natural targets: RTX 3060 12 GB, 4070 12 GB, 5070 12 GB, Mac 16 GB+.
- Gemma 4 12B — Q5_K_M
- ≈ 9.5 GB VRAM. Fits on 12 GB with a modest context; more comfortable on 16 GB (4070 Ti Super, 5080).
- Gemma 4 12B — Q8_0
- ≈ 13 GB VRAM. Reserved for 16 GB+ cards or generous Apple Silicon.
- Gemma 4 26B-A4B — Q4_K_M
- ≈ 19 GB VRAM. Sweet spot: RTX 3090, 4090, 5090, RX 7900 XTX, Mac Studio. MoE: loads only 4B active parameters, so it's fast despite its footprint.
- Gemma 4 26B-A4B — Q5_K_M
- ≈ 23 GB of VRAM. That’s the upper limit for 24 GB; otherwise, target a Mac with 48 GB+ of unified memory or multiple GPUs.
- Gemma 4 26B-A4B — Q8_0
- ≈ 30 GB VRAM. For 32 GB+ (RTX 5090) or Macs with 48 GB or more of unified memory.
#3. Tokens/sec: RTX and Apple Silicon
Typical figures observed with a recent Ollama, 8k context, short prompt, and 200-token generation. The numbers vary by ±15% depending on the generated content and the Ollama version. Since the 26B-A4B is a MoE model (4B active), it runs much faster than a dense 26B once loaded into memory.
- RTX 3060 12 GB
- Gemma 4 E2B: ~90 tok/s. Gemma 4 12B Q4: ~28 tok/s. 26B-A4B: overflows (19 GB), avoid.
- RTX 4070 12 GB
- E2B: ~130 tok/s. 12B Q4: ~46 tok/s. 26B-A4B: overflows.
- RTX 4080 Super 16 GB
- 12B Q4: ~66 tok/s. 12B Q8: ~40 tok/s. 26B-A4B: does not fit in 16 GB.
- RTX 4090 24 GB
- 12B Q4: ~100 tok/s. 26B-A4B Q4: ~58 tok/s. 26B-A4B Q5: ~50 tok/s.
- RTX 5090 32 GB
- 12B Q4: ~150 tok/s. 26B-A4B Q4: ~88 tok/s. 26B-A4B Q8: ~48 tok/s.
- Mac M3 Max 64 GB
- 12B Q4: ~38 tok/s. 26B-A4B Q4: ~30 tok/s. Excellent consistency thanks to unified memory.
- Mac M4 Pro 48 GB
- 12B Q4: ~34 tok/s. 26B-A4B Q4: ~24 tok/s. The MoE runs smoothly and stays responsive for its size.
- Mac M4 Max 128 GB
- 26B-A4B Q4: ~36 tok/s. 26B-A4B Q8: ~20 tok/s. The Mac best suited to using the 26B day to day.
#4. First prompt and useful settings
Gemma 4 performs well in French without a special system prompt, but some parameters are worth adjusting depending on the use case.
To adjust parameters on the fly from the interactive prompt:
- temperature
- 0.7 by default. Lower it to 0.2–0.4 for factual work, code, and RAG. Raise it to 0.9–1.1 for brainstorming.
- num_ctx
- Context window. 4096 or 8192 by default, depending on the builds. Increase it based on your use case and available VRAM.
- num_predict
- Maximum number of tokens generated. -1 for unlimited (until stop). Useful for bounding response length.
- repeat_penalty
- 1.1 by default. If Gemma 4 loops, raise it to 1.15-1.2. If the responses seem too scripted, lower it to 1.05.
#5. API and integration from your code
Ollama exposes an OpenAI-compatible endpoint at http://localhost:11434/v1. Any SDK that speaks OpenAI works by simply pointing base_url to your local instance.
For the multimodal version (12B and 26B-A4B), pass the image as base64 or a local URL in the message—the same protocol as GPT-4o.
#Troubleshooting
- "Error: model gemma4 not found"
- The tag does not yet exist in your version of Ollama, or it is a typo. Use ollama list to check the installed models and the Ollama library documentation for official tags.
- Model spilling over to CPU
- ollama ps affiche 60% GPU / 40% CPU. Passez à une quantization plus agressive (Q4 → Q3_K_S), réduisez num_ctx, ou descendez d'une taille (12B → E2B).
- Speed cut by 3 after a few hours
- The KV cache fragments on some drivers. Restart the Ollama service (sudo systemctl restart ollama on Linux; restart the app on Mac/Windows).
- Responses truncated after ~400 words
- num_predict is too low. Set it to 2048 or -1 for unlimited, and increase num_ctx accordingly.
- Approximate French on E2B
- That's expected: E2B is powerful for its size but remains a compact variant. For demanding French, choose the 12B.
- OOM on RTX 4060 8 GB with 12B
- The 12B Q4 model barely fits in 8 GB, with no room for context. Either use gemma4:e2b-it-qat, accept CPU offload (~3-5 tok/s), or change cards.
#Go further
You have Gemma 4 running. Here are a few ways to take it further, depending on what you want to do with it:
- What is Ollama and how does it work
- If you discover Ollama, this beginner's guide covers the essential commands and the complete mental model.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To understand exactly what you sacrifice by moving from Q8 to Q4, and when it really matters.
- Open WebUI with Ollama
- For a ChatGPT-like interface on top of Gemma 4: conversation history, built-in RAG, multi-user sharing.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.