Intermediate 10 minGemma

Gemma 4 locally with Ollama: installation, VRAM, perfs

Gemma 4 is Google's new generation of open-weight models, released in 2026. Native multimodality, long context, an Apache 2.0 license, and three useful sizes in the Ollama library: E2B (compact), 12B, and 26B-A4B (a MoE). This guide shows you how to run Gemma 4 with Ollama, which quantization to choose based on your VRAM, approximate tokens/sec on RTX and Apple Silicon, and what concretely changes compared with Gemma 3.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run Gemma 4 locally

Gemma 4 is one of the best open-weight models of its size for French. The 12B Q4 fits comfortably in 8 to 12 GB of VRAM, covering a RTX 3060, a 4070, or a midrange M-series Mac. The 26B-A4B, an MoE that activates only 4B parameters, makes use of 24 GB cards (3090, 4090, 7900 XTX) and Macs with plenty of unified memory while remaining fast. For a French-language assistant, RAG, and general-purpose coding, it's an excellent default.

The other reason to run it locally: since April 2026 Gemma 4 has been released under the Apache 2.0 license, lifting the restrictions of the former Gemma Terms of Use (free commercial use, redistribution, and fine-tuning), and Ollama handles pulling, quantization, and inference without requiring you to touch Python, CUDA, or llama.cpp directly.

i
In two words
Gemma 4 + Ollama: one command to pull, one command to chat, and an OpenAI-compatible endpoint at localhost:11434. Everything else is just configuration.

#What changes versus Gemma 3

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Gemma 3 (March 2025) introduced multimodality and a 128k context window. Gemma 4 preserves those gains and pushes forward on three fronts: quality (a major measured leap on reasoning and coding benchmarks), efficiency (the 12B replaces Gemma 3's 27B to good effect on many tasks, and the 26B-A4B MoE runs the equivalent of a large model at the speed of a small one), and tokenizer (a denser vocabulary, which reduces the number of tokens for equivalent output—so more useful tokens per second).

Sizes
Gemma 3: 1B, 4B, 12B, 27B (dense). Gemma 4: E2B (compact), 12B (dense), and 26B-A4B (MoE, 4B active). The lineup scales in efficiency rather than raw size.
Multimodal
Native vision on the 12B and 26B-A4B (text + images). The E2B remains the compact, text-focused variant.
Context
128k tokens across the entire lineup. Same memory constraints as before: the KV cache adds up quickly.
Tokenizer
SentencePiece, with a vocabulary revised to better cover French, code, and non-Latin languages. With equivalent prompts, it uses ~5 to 15% fewer tokens.
License
Apache 2.0 since April 2026. Commercial use, redistribution, and fine-tuning are permitted, with no specific acceptable-use clause—a real change from the Gemma Terms of Use of Gemma 3.
→
When to keep Gemma 3
If you already have a well-tuned RAG pipeline on Gemma 3 12B and the quality is sufficient, do not migrate on principle. Gemma 4 12B is better, but the slight VRAM overhead (~8 GB vs ~7 GB in Q4) and prompt retuning are justified only if you are hitting the limits of Gemma's 12B 3.

#Prerequisites

Ollama installed
Recent version (≥ 0.5). See the Windows, macOS, or Linux installation guides if you haven't already.
Disk space
Allow 4.3 GB for the E2B, 7.6 GB for the 12B Q4, and 19 GB for the 26B-A4B Q4. The Q5 and Q8 variants of the 12B and 26B weigh 1.3 to 2× more.
VRAM or unified memory
Practical minimum: 6 GB for the E2B, 8 to 12 GB for the 12B in Q4, 24 GB for the 26B-A4B in Q4. Below that, it spills onto the CPU and speed collapses.
Up-to-date GPU drivers
CUDA 12.x for NVIDIA, ROCm 6.x for AMD, native Metal for Apple Silicon. Ollama detects the backend automatically.

#1. Pull and launch with one command

The Ollama tag for Gemma 4 follows the usual convention: gemma4 (latest points to the 12B Q4 by default), gemma4:e2b-it-qat, gemma4:12b, gemma4:26b. To specify quantization, add a suffix: gemma4:12b-instruct-q4_K_M.

Direct pull and launch
ollama run gemma4:12b

On the first run, Ollama downloads the model (~7.6 GB for the 12B Q4), then opens an interactive prompt directly. Subsequent launches are instantaneous as long as the model remains in memory cache.

Pull only (without launching)
ollama pull gemma4:e2b-it-qat
ollama pull gemma4:12b
ollama pull gemma4:26b
i
Explicitly choose the quantization
By default, Ollama serves Q4_K_M, which offers the best quality/memory trade-off in 95% of cases. For higher quality: gemma4:12b-instruct-q5_K_M (slightly better quality, +25% VRAM), gemma4:12b-instruct-q8_0 (close to FP16, +100% VRAM). E2B, meanwhile, is distributed directly in QAT int4 (gemma4:e2b-it-qat) and has no useful heavier variant. Raw FP16 is only worthwhile for fine-tuning.

To check what is currently loaded in VRAM:

Loaded model status
ollama ps

The PROCESSOR column must show 100% GPU. If you see 70%/30% GPU/CPU, the model is spilling over—switch to more aggressive quantization or a smaller size.

#2. VRAM by size and quantization

The figures below include the model weights plus a KV cache for an 8k-token context window (the Ollama default). To push to 32k or 128k, budget an additional 2 to 8 GB depending on size.

Gemma 4 E2B — QAT int4
≈ 4.5 GB VRAM. Compact variant (QAT, ~2B active parameters in the MatFormer style). Runs on any GTX 1660 (6 GB), RTX 3050 (8 GB), or Mac with 8 GB of unified memory.
Gemma 4 12B — Q4_K_M
≈ 8 GB of VRAM. Natural targets: RTX 3060 12 GB, 4070 12 GB, 5070 12 GB, Mac 16 GB+.
Gemma 4 12B — Q5_K_M
≈ 9.5 GB VRAM. Fits on 12 GB with a modest context; more comfortable on 16 GB (4070 Ti Super, 5080).
Gemma 4 12B — Q8_0
≈ 13 GB VRAM. Reserved for 16 GB+ cards or generous Apple Silicon.
Gemma 4 26B-A4B — Q4_K_M
≈ 19 GB VRAM. Sweet spot: RTX 3090, 4090, 5090, RX 7900 XTX, Mac Studio. MoE: loads only 4B active parameters, so it's fast despite its footprint.
Gemma 4 26B-A4B — Q5_K_M
≈ 23 GB of VRAM. That’s the upper limit for 24 GB; otherwise, target a Mac with 48 GB+ of unified memory or multiple GPUs.
Gemma 4 26B-A4B — Q8_0
≈ 30 GB VRAM. For 32 GB+ (RTX 5090) or Macs with 48 GB or more of unified memory.
!
The 128k context trap
Enabling the maximum context window (num_ctx=131072) on the 12B can add 4 to 6 GB to the KV cache. If you just fit in Q4 at 8k, you will run out of memory at 32k. Increase the context in increments and monitor ollama ps.

#3. Tokens/sec: RTX and Apple Silicon

Typical figures observed with a recent Ollama, 8k context, short prompt, and 200-token generation. The numbers vary by ±15% depending on the generated content and the Ollama version. Since the 26B-A4B is a MoE model (4B active), it runs much faster than a dense 26B once loaded into memory.

RTX 3060 12 GB
Gemma 4 E2B: ~90 tok/s. Gemma 4 12B Q4: ~28 tok/s. 26B-A4B: overflows (19 GB), avoid.
RTX 4070 12 GB
E2B: ~130 tok/s. 12B Q4: ~46 tok/s. 26B-A4B: overflows.
RTX 4080 Super 16 GB
12B Q4: ~66 tok/s. 12B Q8: ~40 tok/s. 26B-A4B: does not fit in 16 GB.
RTX 4090 24 GB
12B Q4: ~100 tok/s. 26B-A4B Q4: ~58 tok/s. 26B-A4B Q5: ~50 tok/s.
RTX 5090 32 GB
12B Q4: ~150 tok/s. 26B-A4B Q4: ~88 tok/s. 26B-A4B Q8: ~48 tok/s.
Mac M3 Max 64 GB
12B Q4: ~38 tok/s. 26B-A4B Q4: ~30 tok/s. Excellent consistency thanks to unified memory.
Mac M4 Pro 48 GB
12B Q4: ~34 tok/s. 26B-A4B Q4: ~24 tok/s. The MoE runs smoothly and stays responsive for its size.
Mac M4 Max 128 GB
26B-A4B Q4: ~36 tok/s. 26B-A4B Q8: ~20 tok/s. The Mac best suited to using the 26B day to day.
→
Practical reference point
Above 30 tok/s, interactive use feels smooth. Between 15 and 30 tok/s, it's still fine for chat but starts to drag on long outputs. Below 10 tok/s, reserve it for batch jobs or tasks where you would be waiting for completion anyway.

#4. First prompt and useful settings

Gemma 4 performs well in French without a special system prompt, but some parameters are worth adjusting depending on the use case.

Quick test
ollama run gemma4:12b
>>> Explique la différence entre un index B-tree et un index hash en SQL, en 5 lignes.

To adjust parameters on the fly from the interactive prompt:

Session parameters
/set parameter temperature 0.3
/set parameter num_ctx 16384
/set parameter num_predict 800
temperature
0.7 by default. Lower it to 0.2–0.4 for factual work, code, and RAG. Raise it to 0.9–1.1 for brainstorming.
num_ctx
Context window. 4096 or 8192 by default, depending on the builds. Increase it based on your use case and available VRAM.
num_predict
Maximum number of tokens generated. -1 for unlimited (until stop). Useful for bounding response length.
repeat_penalty
1.1 by default. If Gemma 4 loops, raise it to 1.15-1.2. If the responses seem too scripted, lower it to 1.05.
i
Persist settings
To lock in a system prompt and parameters, create a Modelfile: FROM gemma4:12b, then SYSTEM "...", PARAMETER temperature 0.3, etc. ollama create monassistant -f Modelfile, then call monassistant as a normal model.

#5. API and integration from your code

Ollama exposes an OpenAI-compatible endpoint at http://localhost:11434/v1. Any SDK that speaks OpenAI works by simply pointing base_url to your local instance.

Python — openai SDK
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # valeur ignorée, mais le SDK l'exige
)

resp = client.chat.completions.create(
    model="gemma4:12b",
    messages=[
        {"role": "system", "content": "Tu réponds en français, de manière concise."},
        {"role": "user", "content": "Résume la photosynthèse en 3 lignes."},
    ],
    temperature=0.3,
)
print(resp.choices[0].message.content)

For the multimodal version (12B and 26B-A4B), pass the image as base64 or a local URL in the message—the same protocol as GPT-4o.

Vision via curl
curl http://localhost:11434/api/generate -d '{
  "model": "gemma4:12b",
  "prompt": "Décris cette image en français.",
  "images": ["'"$(base64 -w0 photo.jpg)"'"],
  "stream": false
}'

#Troubleshooting

"Error: model gemma4 not found"
The tag does not yet exist in your version of Ollama, or it is a typo. Use ollama list to check the installed models and the Ollama library documentation for official tags.
Model spilling over to CPU
ollama ps affiche 60% GPU / 40% CPU. Passez à une quantization plus agressive (Q4 → Q3_K_S), réduisez num_ctx, ou descendez d'une taille (12B → E2B).
Speed cut by 3 after a few hours
The KV cache fragments on some drivers. Restart the Ollama service (sudo systemctl restart ollama on Linux; restart the app on Mac/Windows).
Responses truncated after ~400 words
num_predict is too low. Set it to 2048 or -1 for unlimited, and increase num_ctx accordingly.
Approximate French on E2B
That's expected: E2B is powerful for its size but remains a compact variant. For demanding French, choose the 12B.
OOM on RTX 4060 8 GB with 12B
The 12B Q4 model barely fits in 8 GB, with no room for context. Either use gemma4:e2b-it-qat, accept CPU offload (~3-5 tok/s), or change cards.

#Go further

You have Gemma 4 running. Here are a few ways to take it further, depending on what you want to do with it:

What is Ollama and how does it work
If you discover Ollama, this beginner's guide covers the essential commands and the complete mental model.
Choose your quantization (Q4, Q5, Q8, FP16)
To understand exactly what you sacrifice by moving from Q8 to Q4, and when it really matters.
Open WebUI with Ollama
For a ChatGPT-like interface on top of Gemma 4: conversation history, built-in RAG, multi-user sharing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.