Beginner 9 minOllama

Gemma 3 (Google) locally: complete guide 2026

Gemma 3 is Google's open-weight model family released in March 2025. Four sizes (1B, 4B, 12B, 27B), vision integrated starting at 4B, 128k context, and solid multilingual support that includes decent French. This guide covers installing Gemma 3 Ollama locally step by step, the Google license you need to know, and where the model really shines.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Gemma 3?

Gemma 3 is the third iteration of Google DeepMind's open-weight models, derived from the same research as Gemini. The family covers four very distinct sizes — from 1B to 27B — allowing you to target exactly the hardware tier you have available, from a laptop without a GPU to a 24 GB workstation.

Three things set Gemma 3 apart from other models in this category: a 128k context window by default (only 32k for the 1B), integrated multimodal vision starting with the 4B without having to load a separate model, and explicit work on multilingual support (Google’s documentation claims coverage of 140+ languages for the 4B sizes and above).

i
In two words
Gemma 3 = Google's “versatile” family. Not the strongest on any particular benchmark, but one of the most complete: 4 sizes, vision, long context, multilingual. Ideal as the default model on a given machine.

#The 4 Gemma sizes 3

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Choosing the right size is the most important decision. Here is the usage-to-hardware mapping:

gemma3:1b
Text only, 32k context. For pure CPU use, Raspberry Pi 5, or an 8 GB Mac M1. Limited quality but instant responses. Good for classification and short rewriting.
gemma3:4b
Multimodal (text + image), 128k context. The “laptop” tier: fits comfortably in 6 GB of VRAM or on a 16 GB M2/M3 Mac. Reasonable quality for general-purpose tasks.
gemma3:12b
Multimodal, 128k context. The “mid-range PC” tier: RTX 3060 12 GB, 4070, M3 Pro 18 GB. A quality level that's genuinely usable every day.
gemma3:27b
Multimodal, 128k context. The “workstation” tier: RTX 3090/4090, M3 Max. The best Gemma 3, comparable to higher-tier models on many tasks.
→
If you're unsure
Start with the 4B. It’s the smallest model that includes vision and a 128k context, and it runs practically everywhere. Move up to 12B or 27B once you know whether the quality is good enough for you.

#VRAM requirements (Q4_K_M quantization)

gemma3:1b
≈ 1 GB of VRAM or RAM. Runs on any machine, including a Raspberry Pi 5.
gemma3:4b
≈ 3 GB of VRAM. RTX 3050 6 GB, GTX 1660 6 GB, and an 8 GB MacBook Air M2 are sufficient.
gemma3:12b
≈ 8 GB of VRAM. RTX 3060 12 GB, RTX 4070 12 GB, Mac M3 Pro 18 GB unified.
gemma3:27b
≈ 17 GB of VRAM in Q4. RTX 3090/4090 24 GB or Mac M3/M4 Max 36 GB+. In Q3 (degraded quality), it runs on 16 GB.

Add about 1 to 4 GB for the KV cache, depending on the context length you actually use. Loading the full 128k context of a 27B model requires more than the model’s raw VRAM.

#1. Local installation of Gemma 3 Ollama

If Ollama is not installed yet, first follow the Ollama installation guide for your OS. Once Ollama is running, downloading Gemma 3 takes one command.

Terminal
ollama run gemma3:4b

Ollama downloads the model (about 3.3 GB for the 4B in Q4), then launches it directly in interactive mode. When the >>> prompt appears, you can test it.

Pull without launching
ollama pull gemma3:1b
ollama pull gemma3:4b
ollama pull gemma3:12b
ollama pull gemma3:27b

The Ollama daemon listens on http://localhost:11434 — you can connect Open WebUI, LM Studio as clients, Continue.dev, or any OpenAI-compatible client to it.

→
Choosing the right quantization
The default tag (e.g., gemma3:12b) points to Q4_K_M, the best quality/memory compromise for most use cases. For higher quality at the cost of more VRAM, use gemma3:12b-it-q5_K_M or q8_0.

#2. Multilingual and French quality

Google's work on multilingual support remains a real strength of Gemma 3 compared with smaller Anglocentric models. French is well represented in the training corpus — you will almost never see the model switch to English mid-response, which remains a recurring flaw in smaller models. The 2026 generations (Qwen 3.5, Granite 4.2) have caught up significantly in French, but Gemma 3 still holds its own in this area.

Gemma 3 4B
Correct French for simple tasks (summarization, rephrasing, classification). For long-form writing, the phrasing is sometimes awkward.
Gemma 3 12B
“Usable French assistant” level. Good command of grammar, varied vocabulary, and few distracting anglicisms. Comparable to Qwen 3.5 9B on this criterion.
Gemma 3 27B
Very good French. On par with Mistral Small 24B for writing, slightly behind on formal reasoning but often better at nuance and style.
Quick test in French
>>> Résume en 3 points clés les enjeux de la souveraineté numérique européenne.

1. **Dépendance technologique** : 80% du cloud européen repose sur trois acteurs américains...
2. **Conformité réglementaire** : le RGPD et l'AI Act créent un cadre que les fournisseurs hors-UE...
3. **Capacité industrielle** : la course aux semi-conducteurs et aux modèles d'IA...

#3. Integrated multimodal vision

Starting with 4B, Gemma 3 accepts images as input — no need to load a separate vision model. The same binary handles text and vision, with an integrated SigLIP encoder. Useful for OCR, image descriptions, or screenshot analysis.

Ollama CLI with image
ollama run gemma3:12b
>>> Décris cette capture d'écran et identifie les éléments d'interface : /chemin/vers/screenshot.png

In the Python API, pass the image as base64 or via a local path:

Ollama Python API
import ollama

response = ollama.chat(
    model='gemma3:12b',
    messages=[{
        'role': 'user',
        'content': 'Extrais le texte visible sur cette facture.',
        'images': ['/chemin/vers/facture.jpg'],
    }]
)
print(response['message']['content'])
i
Realistic vision quality
Gemma 3 Vision is adequate for general descriptions, OCR of clean documents, and reading interfaces. For OCR on blurry or handwritten photos, a newer, larger vision model such as Qwen 3.8 27B (Vision) is still better. For complex spatial reasoning (counting objects, understanding a technical diagram), everything remains imperfect—that is true of all open-weight LLMs.

#4. 128k context in practice

The 128k context of Gemma 3 (except for the 1B model, limited to 32k) is sufficient to ingest a short book, a complete documentation set, or several hours of transcription. By default, Ollama loads a much shorter context—typically 4k or 8k—to save VRAM—you must extend it explicitly.

Extend the context in the CLI
ollama run gemma3:12b
>>> /set parameter num_ctx 32768

Or through the API, passing the num_ctx option when making the request:

Extended context via API
import ollama

response = ollama.chat(
    model='gemma3:12b',
    messages=[{'role': 'user', 'content': 'Résume ce document : ...'}],
    options={'num_ctx': 32768},
)
!
Context costs VRAM
Each context token uses memory for the KV cache. Going from 8k to 128k on a 12B model can add 4–6 GB of consumed VRAM. If you max out, the model switches to CPU offload and speed collapses. Increase gradually: 16k → 32k → 64k based on your actual needs.

#Gemma license: what you need to know

Gemma 3 is NOT licensed under Apache 2.0 or MIT like Mistral, Qwen, or DeepSeek. Google uses its own license: the "Gemma Terms of Use." It allows commercial use, including distribution of derivative models, but imposes two constraints that genuinely free licenses do not.

Acceptable use policy
You must comply with the Gemma Prohibited Use Policy (no CSAM generation, no blatant rights violations, etc.). The list may be revised unilaterally by Google.
Term propagation
If you redistribute the model (or a derivative), you must provide the Gemma license to your users and require them to accept the same restrictions.
No “copyleft”
Your applications that call Gemma 3 are NOT contaminated by the license. Only redistributed weights are covered.
!
For professional use
For simple use as an internal SaaS or product (the end user never sees the weights), Gemma 3 is usable without issue. To publish a fine-tune or integrate Gemma 3 into an open-source product distributed with the weights, read the license first—the difference from a truly permissive license becomes significant. Note: Google moved the next generation, Gemma 4, to the Apache 2.0 license (April 2026)—if weight redistribution matters to you, it removes precisely this constraint.

#Gemma 3 27B vs Llama 4 Scout

Both models target the “accessible high-end for local use” segment. To choose, you need to compare what really sets them apart:

Architecture
Gemma 3 27B is a dense model (27B parameters, all active). Llama 4 Scout is a MoE model (17B active, 109B total) — much more capable at equivalent VRAM when it fits, but it requires loading all the experts.
Minimum VRAM
Gemma 3 27B Q4: ≈ 17 GB. Llama 4 Scout Q4: ≈ 60 GB (all experts loaded). On a single RTX 4090, only Gemma 3 27B fits comfortably.
Context
Gemma 3: 128k. Llama 4 Scout: 10M (theoretical, more modest in practice). If you’re actually processing more than 128k tokens, Scout is ahead.
Vision
Both are natively multimodal. Comparable quality on general-purpose tasks.
FR
Gemma 3 is slightly better in everyday French. Llama 4 Scout is stronger at formal reasoning and code.
License
Gemma Terms of Use vs. Llama 4 Community License. Both are non-OSI, and both are commercially acceptable with specific restrictions.
→
Verdict
On a single 24 GB GPU: Gemma 3 27B. On a Mac Studio Ultra or a multi-GPU system: Llama 4 Scout becomes relevant for the benefits of MoE. Many users keep both installed and switch depending on the task.

#Troubleshooting

"Out of memory" while loading
Insufficient VRAM for the selected size. Drop down one size (gemma3:12b → gemma3:4b) or use more aggressive quantization (Q4 → Q3). Do not force CPU offload on a 27B model: it becomes unusable.
Rejected images
You're probably using gemma3:1b, which is text-only. Upgrade to at least gemma3:4b for vision support.
Context ignored beyond 4k
Ollama loads a very short num_ctx by default. Extend it explicitly with /set parameter num_ctx 32768 or via the API option options={'num_ctx': 32768}.
Slow generation on 12B/27B with a GPU
Check with ollama ps that the PROCESSOR column correctly shows 100% GPU. If you see a CPU percentage, VRAM plus context do not fit—reduce num_ctx.
Truncated responses
Increase num_predict (128 by default in some clients). For Ollama CLI: /set parameter num_predict 2048.

#Go further

Gemma 3 is an excellent all-purpose starting point. Here are some natural directions depending on what you want to use it for:

Compare with Gemma 4
The guide “Gemma 4 locally with Ollama” shows how the family has evolved and where upgrading is worthwhile.
Choosing the right quantization
The “Choosing Your Quantization (Q4, Q5, Q8, FP16)” guide details the trade-offs that apply particularly well to the 4 sizes Gemma 3.
Using vision in a pipeline
The “Local multimodal vision LLM” guide provides concrete Python recipes (OCR, description, analysis) directly applicable to Gemma 3 4B/12B/27B.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.