Intermediate 11 minGemma

Migrating from Gemma 2 to Gemma 3: pitfalls and checks

Direct response

Moving from Gemma 2 to Gemma 3 rarely breaks model startup, but often breaks behavior: the context grows from 8,192 tokens to 32,768 (1B) or 131,072 tokens (4B/12B/27B), the chat template changes, and 4B/12B/27B become multimodal while Gemma 2 was text-only. Recheck the template applied by your runtime, the actually configured context, and rerun your test prompts before any production switch.

Gemma 2 (9B, 27B) and Gemma 3 (1B, 4B, 12B, 27B) are not the same family with an extra version number: the architecture, context, and input format change. This guide lists the issues that actually break an application already built on Gemma 2, the checks to perform before replacing the model in production, and what you need to know to roll back if necessary.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What really changes between Gemma 2 and Gemma 3

Three structural changes affect an existing application: the maximum context, the structure of the messages sent to the model, and whether an image input is present. Simply changing the model name in your code (gemma2 to gemma3) isn't enough to guarantee identical behavior, even if the output format remains text.

Gemma 2 existed only in 9 and 27 billion-parameter versions. Gemma 3 adds a 1B and a 4B, and redefines the context for each size: it is not merely “larger”—the internal architecture is different. If your model-size choice was based on the constraints of Gemma 2, it should be reconsidered rather than carried over unchanged.

#Size and context: the real leap

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Gemma 2 has a fixed context of 8,192 tokens, regardless of model size. Gemma 3 raises this to 32,768 tokens for the 1B and to 131,072 tokens for the 4B, 12B, and 27B models: a factor of 16 for medium and large sizes. This is the most useful data point for sizing a RAG workload or a long conversation history, but the advertised context is a model limit, not the memory actually allocated by your runtime: Ollama and inference servers often limit the context by default to a much lower value (often 4096 or 8192 tokens) until you explicitly increase it.

Gemma 2 vs Gemma 3: what changed by size
SizeGemma 2 contextGemma 3 contextInput image
1B—32 768 tokensNo (text only)
4B—131,072 tokensYes
9B / 12B8,192 tokens (9B)131,072 tokens (12B)9B: no · 12B: yes
27B8,192 tokens131,072 tokensYes
i
Surprise gradient
Increasing the context declared in your configuration isn't enough: the KV cache grows with the context actually used, not the theoretical maximum. A 131,072-token context opened by default can saturate GPU memory before the first generation if the runtime preallocates this cache.

#Chat template: the most common pitfall

The most common cause of degraded responses after a migration is not model quality, but an incorrectly applied chat template. Each model family expects precise formatting for conversation turns (turn start/end tokens, system prompt position). If your application builds the final prompt itself instead of letting the runtime apply the model’s template, a template still set to Gemma 2 produces truncated, irrelevant responses or continues generating past the expected end.

  1. 01
    Identify who builds the prompt
    Check whether your code assembles the text sent to the model itself or goes through the runtime's chat API (Ollama /api/chat, llama.cpp server), which automatically applies the correct template.
  2. 02
    Compare the templates
    Retrieve the template file embedded in the Gemma 3 model being used (GGUF or Hugging Face) and compare it with the one used by Gemma 2 so far, rather than assuming they are identical.
  3. 03
    Replay a reference prompt set
    Run the same 10 to 20 test prompts on Gemma 2 and then Gemma 3 with the new template, and compare length, relevance, and whether generation ends cleanly.

The system role illustrates this pitfall well. When presenting Gemma 4, Google states that this new family “uses standard system, assistant, and user roles, unlike Gemma 3”: official confirmation that Gemma 3 (like Gemma 2) does not handle the system prompt exactly like runtimes that expose a native system role. In practice, code that sends a separate system-role message based on conventions from other model families must check, specifically on Gemma 3, how its runtime handles that message rather than assuming it is isolated.

#Multimodal: 4B, 12B, and 27B accept images

Unlike the entirely text-only Gemma 2, the Gemma 3 4B, 12B, and 27B models include a SigLIP image encoder that processes images resized to 896×896 pixels; only the 1B remains text-only. In practice, if your application migrates to a 12B or 27B model, it inherits image-input capability even if it does not use it: this changes the API format expected by some runtimes (an image field in addition to text) and can slightly increase the memory required for loading, even when no image is sent.

This shift to multimodality has a practical consequence that is often overlooked during migration: a pipeline that strictly validated the format of incoming messages (JSON schema, strict typing) may reject or misinterpret an absent or empty image field if the new model's API client adds it by default. Conversely, a pipeline that already filtered non-text inputs before calling the model needs no changes: the image capability of Gemma 3 remains inactive until an image is sent; it is never required.

#QAT quantization: memory reduced by 3 to 4 times

Google publishes Gemma 3 checkpoints trained with quantization awareness (QAT), in addition to the standard Q4_0 versions for Ollama, llama.cpp, and MLX. The benefit: significantly less quality loss than quantizing afterward a model that was not prepared for it.

Memory reported by Google, BF16 vs QAT int4
SizeBF16QAT int4
27B54 GB14.1 GB
12B24 GB6.6 GB
4B8 GB2.6 GB
1B2 GB0.5 GB
→
What this changes for the hardware
A 27B QAT in int4 fits on a 24 GB VRAM card according to the manufacturer's figures, which wasn't true of a Gemma 2 27B at native precision. If your migration also involves changing model size, check the memory actually available on your machine before choosing the target size, not just the advertised BF16 figure.

These figures were published by Google when the QAT checkpoints launched; they describe the memory required to load the weights, not the total memory used during generation (the context and KV cache add to it). Measuring on your own machine is still the only way to confirm a figure for your use case.

Google's method applies approximately 5,000 steps of quantization-aware training using the unquantized model's probabilities as the target, reducing the perplexity drop by 54% compared with post-training quantization applied to a model that was not prepared for it. For an application migrated from Gemma 2, this means that the same quantization level (for example, Q4) generally produces a result on Gemma 3 that is closer to the full-precision model than what conventionally quantized Gemma 2 produced—a gain that comes from preparing the model, not from a setting the user needs to adjust.

#Migrate to Gemma 3 or wait for Gemma 4?

Gemma 4 is already available on Ollama when this guide is written, raising a real question for anyone still migrating from Gemma 2: should you stop at Gemma 3 or target the next generation directly? The Ollama specification for Gemma 4 lists sizes different from those of Gemma 3 (E2B and E4B variants with reduced effective parameters, a 12B, a 26B MoE variant with 3.8 billion active parameters, and a 31B dense model), positioned for reasoning, agentic workflows, and code—a broader scope than simply replacing Gemma 2.

For an application already in production under Gemma 2, this guide remains focused on migrating to Gemma 3: the context, template, and regression checks described here apply identically if the final target is Gemma 4, but the direct switch changes more things (standard chat roles, new sizes, MoE architecture on the 26B variant) than a Gemma 2 → Gemma 3 migration. A dedicated guide covers installing and running Gemma 4 locally.

i
What Gemma 4 changes in practice
The system role becomes native (system/assistant/user) instead of being handled in the style of Gemma 2/3, which simplifies integration in application code but requires revalidating the message format a second time if you jump directly to Gemma 4 instead of stopping at Gemma 3.

#Checklist before moving to production

Runtime version
Check that Ollama, llama.cpp, or MLX support the targeted version 3 of Gemma (support was added after the model was released and is not always immediate in an older runtime version).
Actually configured context
Don't leave the server's context value to chance: set it explicitly based on your actual needs, not the model's ceiling.
Image input format
If you migrate to 4B, 12B, or 27B, check that your API client isn't sending an image format incompatible with the new endpoint.
Quantization selection
Compare a classic GGUF Q4_K_M and an official QAT checkpoint on your own prompts before locking in a choice; both are available for Gemma 3.
Rollback window
Keep model Gemma 2 and its configuration available during the transition, until you confirm there is no regression.

#Detect a regression in your prompts

You can’t validate a model migration based on how it feels after two or three questions. A reference prompt set that represents the application’s real-world use, replayed identically against the old and new models, is the only way to detect a silent regression: a longer but less precise response, loss of an expected output format (JSON, list), or tone drift. Keeping these reference outputs also lets you document the migration decision instead of relying on an impression.

!
Common objection: “the docs say it’s compatible”
API-level compatibility (same endpoint, same request format) is not behavioral compatibility. The context, template, and image input genuinely differ between the two families; only a test with your own prompts can confirm that no regression has slipped in.

#Plan a rollback

The safest scenario for a production application is to deploy Gemma 3 in parallel with Gemma 2 (a distinct model name in the runtime), route some traffic to the new version, and shut down the old one only when quality measurements on your reference prompts are at least equivalent. This costs some disk space and RAM during the transition, but avoids a service outage if the new chat template or context produces unexpected behavior in real-world conditions.

Frequently asked questions
Can you replace Gemma 2 with Gemma 3 without changing the code?+
The model name changes without breaking the API call in most runtimes, but behavior may change: context expanded up to 16×, a different chat template, and image input available on 4B/12B/27B. A test with your own prompts, compared with the output from Gemma 2, is still necessary before any production switch.
Is Gemma 3 more demanding to run than Gemma 2?+
At a comparable size for the 27B, no thanks to the QAT checkpoints: Google reports 14.1 GB of VRAM in int4 versus 54 GB in BF16, enough to fit on a single RTX 3090 24 GB according to the manufacturer. But actually using a larger context increases the KV cache, so total memory consumption also increases during generation.
Should you use the QAT version or a standard Q4_K_M GGUF?+
Both exist for Gemma 3. The QAT version is trained with approximately 5,000 steps of quantization-aware training to limit quality loss, reducing the perplexity increase by 54% compared with quantization applied afterward. A classic Q4_K_M remains an option if your runtime does not yet support official QAT checkpoints.
Is Gemma 3’s 131,072-token context enabled by default?+
No. That's the model's ceiling, and the Ollama spec sheet for Gemma 3 states a default context of 128K rather than automatically enabling the theoretical maximum. Most runtimes limit the actually allocated context to a much lower value until it is explicitly increased in the server configuration.
Should you migrate directly to Gemma 4 rather than Gemma 3?+
It depends on your needs: Gemma 4 adds standard chat roles (system/assistant/user) and different sizes (E2B, E4B, 12B, a 26B MoE variant, and a dense 31B), focused on reasoning and coding. A direct migration changes more reference points than moving through Gemma 3, so it deserves its own regression test rather than a blind switch.
Will a Gemma 3 27B model fit on a consumer graphics card?+
Yes for the int4 QAT version: Google states that it fits on a RTX 3090 with 24 GB of VRAM, compared with the 54 GB required by native BF16. This figure covers loading the weights, not the context or KV cache used during generation, which are added according to the actual length of the exchanges.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.