Gemma 3 (Google) locally: complete guide 2026
Gemma 3 is Google's open-weight model family released in March 2025. Four sizes (1B, 4B, 12B, 27B), vision integrated starting at 4B, 128k context, and solid multilingual support that includes decent French. This guide covers installing Gemma 3 Ollama locally step by step, the Google license you need to know, and where the model really shines.
#Why Gemma 3?
Gemma 3 is the third iteration of Google DeepMind's open-weight models, derived from the same research as Gemini. The family covers four very distinct sizes — from 1B to 27B — allowing you to target exactly the hardware tier you have available, from a laptop without a GPU to a 24 GB workstation.
Three things set Gemma 3 apart from other models in this category: a 128k context window by default (only 32k for the 1B), integrated multimodal vision starting with the 4B without having to load a separate model, and explicit work on multilingual support (Google’s documentation claims coverage of 140+ languages for the 4B sizes and above).
#The 4 Gemma sizes 3
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Choosing the right size is the most important decision. Here is the usage-to-hardware mapping:
- gemma3:1b
- Text only, 32k context. For pure CPU use, Raspberry Pi 5, or an 8 GB Mac M1. Limited quality but instant responses. Good for classification and short rewriting.
- gemma3:4b
- Multimodal (text + image), 128k context. The “laptop” tier: fits comfortably in 6 GB of VRAM or on a 16 GB M2/M3 Mac. Reasonable quality for general-purpose tasks.
- gemma3:12b
- Multimodal, 128k context. The “mid-range PC” tier: RTX 3060 12 GB, 4070, M3 Pro 18 GB. A quality level that's genuinely usable every day.
- gemma3:27b
- Multimodal, 128k context. The “workstation” tier: RTX 3090/4090, M3 Max. The best Gemma 3, comparable to higher-tier models on many tasks.
#VRAM requirements (Q4_K_M quantization)
- gemma3:1b
- ≈ 1 GB of VRAM or RAM. Runs on any machine, including a Raspberry Pi 5.
- gemma3:4b
- ≈ 3 GB of VRAM. RTX 3050 6 GB, GTX 1660 6 GB, and an 8 GB MacBook Air M2 are sufficient.
- gemma3:12b
- ≈ 8 GB of VRAM. RTX 3060 12 GB, RTX 4070 12 GB, Mac M3 Pro 18 GB unified.
- gemma3:27b
- ≈ 17 GB of VRAM in Q4. RTX 3090/4090 24 GB or Mac M3/M4 Max 36 GB+. In Q3 (degraded quality), it runs on 16 GB.
Add about 1 to 4 GB for the KV cache, depending on the context length you actually use. Loading the full 128k context of a 27B model requires more than the model’s raw VRAM.
#1. Local installation of Gemma 3 Ollama
If Ollama is not installed yet, first follow the Ollama installation guide for your OS. Once Ollama is running, downloading Gemma 3 takes one command.
Ollama downloads the model (about 3.3 GB for the 4B in Q4), then launches it directly in interactive mode. When the >>> prompt appears, you can test it.
The Ollama daemon listens on http://localhost:11434 — you can connect Open WebUI, LM Studio as clients, Continue.dev, or any OpenAI-compatible client to it.
#2. Multilingual and French quality
Google's work on multilingual support remains a real strength of Gemma 3 compared with smaller Anglocentric models. French is well represented in the training corpus — you will almost never see the model switch to English mid-response, which remains a recurring flaw in smaller models. The 2026 generations (Qwen 3.5, Granite 4.2) have caught up significantly in French, but Gemma 3 still holds its own in this area.
- Gemma 3 4B
- Correct French for simple tasks (summarization, rephrasing, classification). For long-form writing, the phrasing is sometimes awkward.
- Gemma 3 12B
- “Usable French assistant” level. Good command of grammar, varied vocabulary, and few distracting anglicisms. Comparable to Qwen 3.5 9B on this criterion.
- Gemma 3 27B
- Very good French. On par with Mistral Small 24B for writing, slightly behind on formal reasoning but often better at nuance and style.
#3. Integrated multimodal vision
Starting with 4B, Gemma 3 accepts images as input — no need to load a separate vision model. The same binary handles text and vision, with an integrated SigLIP encoder. Useful for OCR, image descriptions, or screenshot analysis.
In the Python API, pass the image as base64 or via a local path:
#4. 128k context in practice
The 128k context of Gemma 3 (except for the 1B model, limited to 32k) is sufficient to ingest a short book, a complete documentation set, or several hours of transcription. By default, Ollama loads a much shorter context—typically 4k or 8k—to save VRAM—you must extend it explicitly.
Or through the API, passing the num_ctx option when making the request:
#Gemma license: what you need to know
Gemma 3 is NOT licensed under Apache 2.0 or MIT like Mistral, Qwen, or DeepSeek. Google uses its own license: the "Gemma Terms of Use." It allows commercial use, including distribution of derivative models, but imposes two constraints that genuinely free licenses do not.
- Acceptable use policy
- You must comply with the Gemma Prohibited Use Policy (no CSAM generation, no blatant rights violations, etc.). The list may be revised unilaterally by Google.
- Term propagation
- If you redistribute the model (or a derivative), you must provide the Gemma license to your users and require them to accept the same restrictions.
- No “copyleft”
- Your applications that call Gemma 3 are NOT contaminated by the license. Only redistributed weights are covered.
#Gemma 3 27B vs Llama 4 Scout
Both models target the “accessible high-end for local use” segment. To choose, you need to compare what really sets them apart:
- Architecture
- Gemma 3 27B is a dense model (27B parameters, all active). Llama 4 Scout is a MoE model (17B active, 109B total) — much more capable at equivalent VRAM when it fits, but it requires loading all the experts.
- Minimum VRAM
- Gemma 3 27B Q4: ≈ 17 GB. Llama 4 Scout Q4: ≈ 60 GB (all experts loaded). On a single RTX 4090, only Gemma 3 27B fits comfortably.
- Context
- Gemma 3: 128k. Llama 4 Scout: 10M (theoretical, more modest in practice). If you’re actually processing more than 128k tokens, Scout is ahead.
- Vision
- Both are natively multimodal. Comparable quality on general-purpose tasks.
- FR
- Gemma 3 is slightly better in everyday French. Llama 4 Scout is stronger at formal reasoning and code.
- License
- Gemma Terms of Use vs. Llama 4 Community License. Both are non-OSI, and both are commercially acceptable with specific restrictions.
#Troubleshooting
- "Out of memory" while loading
- Insufficient VRAM for the selected size. Drop down one size (gemma3:12b → gemma3:4b) or use more aggressive quantization (Q4 → Q3). Do not force CPU offload on a 27B model: it becomes unusable.
- Rejected images
- You're probably using gemma3:1b, which is text-only. Upgrade to at least gemma3:4b for vision support.
- Context ignored beyond 4k
- Ollama loads a very short num_ctx by default. Extend it explicitly with /set parameter num_ctx 32768 or via the API option options={'num_ctx': 32768}.
- Slow generation on 12B/27B with a GPU
- Check with ollama ps that the PROCESSOR column correctly shows 100% GPU. If you see a CPU percentage, VRAM plus context do not fit—reduce num_ctx.
- Truncated responses
- Increase num_predict (128 by default in some clients). For Ollama CLI: /set parameter num_predict 2048.
#Go further
Gemma 3 is an excellent all-purpose starting point. Here are some natural directions depending on what you want to use it for:
- Compare with Gemma 4
- The guide “Gemma 4 locally with Ollama” shows how the family has evolved and where upgrading is worthwhile.
- Choosing the right quantization
- The “Choosing Your Quantization (Q4, Q5, Q8, FP16)” guide details the trade-offs that apply particularly well to the 4 sizes Gemma 3.
- Using vision in a pipeline
- The “Local multimodal vision LLM” guide provides concrete Python recipes (OCR, description, analysis) directly applicable to Gemma 3 4B/12B/27B.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.