Qwen3 GGUF: fix the tokenizer and chat template
Tokenizer and chat-template errors in Qwen3 GGUFs appear as three symptoms: a reasoning tag that never closes, generation that never stops, or irrelevant answers. The cause is almost always an incorrectly applied chat template. Add --jinja to use the template embedded in the GGUF, verify the file's provenance, and as a last resort pass a custom template.
A GGUF Qwen3 that “answers nonsense” is almost never a model quality issue: it’s a prompt formatting issue before the prompt reaches the model. This guide covers only Qwen3 GGUF tokenizer and chat template errors: recognizing the symptom, checking the file’s provenance, and fixing or replacing the template.
#The three symptoms to recognize
Three signals point to a tokenizer or chat-template problem rather than a model problem: a thinking tag (usually “think”) that opens and never closes in the response, generation that continues indefinitely without stopping at the logical end of the response, or a model that responds off-topic as if it did not understand that it was being asked a question. In all three cases, the model itself is not at fault: the structure of the text it receives as input is malformed.
#The root cause: the chat template
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A language model never receives your messages directly: a chat template formats them (turn tags, system prompt, start/end markers) before transforming them into tokens. If this template is missing, incorrectly chosen, or misinterpreted by the inference engine, the model receives text that does not resemble what it was trained on and produces degraded outputs, even if the model weights and tokenizer themselves are correct.
#The --jinja option: the first thing to check
The official Qwen documentation explicitly recommends adding --jinja when launching a Qwen3 GGUF with llama.cpp: this option tells it to use the chat template embedded in the GGUF file, presented as the preferred method over a generic template selected by the engine by default.
If your launch command does not contain --jinja, this is the first fix to try before considering any other hypothesis. Many scripts and interfaces built before this option became widespread still omit it, which explains a large share of reports of “broken responses” from otherwise valid Qwen3 GGUFs.
#A known parsing bug, now fixed
A specific error affected llama.cpp's Qwen3 chat template: the template engine failed to parse a Jinja list-slicing syntax (messages[::-1]) used to traverse the conversation history backward in the tool-calling logic. The reported error was a parsing failure pointing precisely to that line of the template.
- 01Identify the llama.cpp version in useAn older version may not include the parsing fix for the slicing syntax used by the Qwen3 template.
- 02Update to a recent versionRecompiling or redownloading an up-to-date llama.cpp binary resolves this specific case without modifying the GGUF itself.
- 03If updating isn't possiblePass a simplified custom template via --chat-template-file, avoiding the problematic Jinja construction.
#llama-server and llama-cli do not behave the same way
A behavior reported in the llama.cpp repository: enabling --jinja with llama-server can make the reasoning block—the content between the thinking tags—disappear from the response, while the same block remains visible when using llama-cli with the same option and model. If your integration depends on reasoning content being present in the output, for observability or debugging, this isn't a tokenizer issue but a processing difference between the two binaries—check which one you're using before investigating further.
This distinction has a practical consequence for anyone building an integration on top of llama.cpp rather than a simple interactive session: a test pipeline that validates the output format with llama-cli and is then deployed in production behind llama-server may show a silent regression on this specific point even though no application-side parameter has changed. Explicitly documenting which binary serves production, and testing with that exact binary rather than the one used in local development, avoids this trap.
#Disable forced reasoning mode
Qwen3 provides a mechanism for switching between thinking and direct modes at the chat template level. However, the official documentation Qwen indicates that this forced deactivation mechanism (hard switch) is not natively exposed in llama.cpp: passing enable_thinking as false through the command-line options may be ignored depending on the version, as shown by several recent reports involving Qwen3.5 variants.
The workaround documented by Qwen is to provide a custom template via --chat-template-file, with enable_thinking explicitly set to false at the template level itself rather than passed as a parameter at request time. This is more reliable than a runtime parameter that depends on the exact support in your version of llama.cpp.
One point needs clarification to avoid overgeneralizing: report #20182 ("enable_thinking param cannot turn off thinking") specifically concerns Qwen3.5-9B in build 8215, remains labeled "bug-unconfirmed" in the llama.cpp repository, and was closed without resolution ("not planned"). There is no evidence that the same behavior affects an original Qwen3 GGUF (as opposed to Qwen3.5): if you encounter this symptom on a standard Qwen3, treat it as a case to isolate and report separately rather than as automatic confirmation of this ticket.
#Verify the provenance of a third-party GGUF
Some tokenizer issues with Qwen3 GGUF files do not come from llama.cpp but from the GGUF file itself: a conversion made with an old version of the conversion tools, or a file whose tokenizer was exported incorrectly, produces similar symptoms (missing end of generation, special tokens not recognized correctly). Before looking for a bug in the inference engine, comparing your GGUF file's size and publication date with a recognized repository (Qwen official, or documented requantizations) can rule out this hypothesis.
A simple way to quickly distinguish a file problem from a configuration problem: if the same GGUF works correctly on another machine, or with another version of llama.cpp, the file itself is probably not at fault. Conversely, if multiple re-downloads of the same repository on the same machine consistently reproduce the symptom, the most likely explanation shifts toward the local configuration (binary version, launch options) rather than the file.
#Quantization that is too low: malformed tool calls
A final symptom, distinct from the first three, specifically affects using Qwen3 as an agent with tool calls: instead of a text-formatting problem, the tool call itself arrives truncated, with empty or malformed arguments (invalid JSON). A community troubleshooting guide dedicated to llama.cpp documents that tool-call structure is sensitive to the quantization level: sub-4-bit quantizations (Q3, Q2, IQ) produce malformed tool calls even when the chat template is correctly applied with --jinja.
This issue is easy to miss because, at first glance, it looks like a classic tokenizer problem: a truncated response immediately suggests an improperly closed chat template. The practical distinction lies in the context where the symptom appears: a template problem affects every response, including plain text without a tool call, whereas a quantization problem affecting tool calls generally leaves free-form text responses untouched and appears only in the strict JSON structure expected by the tool-calling protocol.
#Quick troubleshooting table
| Symptom | Most likely cause | Correction first |
|---|---|---|
| Never-closed « think » tag | Chat template not applied | Add --jinja when launching |
| Generation that never stops | Template misinterpreted or missing | Check --jinja; otherwise update llama.cpp |
| “Expected value expression” error on launch | Bug in Jinja slicing parsing (fixed by PR #13573) | Update to a recent version of llama.cpp |
| Reasoning block absent with llama-server but visible with llama-cli | Documented processing difference between the two binaries | Test with llama-cli to confirm, then follow ticket #14894 |
| enable_thinking=false ignored | Hard switch not natively exposed in llama.cpp | Set enable_thinking=false in a template via --chat-template-file |
| Truncated tool calls or invalid JSON | Quantization too low (Q3, Q2, IQ) | Move up to at least Q4_K_M, ideally Q5_K_M or Q6_K |
| Newer GGUF more buggy than an older one for the same model | Poorly converted file or corrupted download | Download again from the original source (official Qwen or recognized repository) |
- Understanding GGUF and safetensors formats
- Qwen3-32B technical specification
- llama.cpp: what is it, and should you leave Ollama?
- Choose your quantization (Q4, Q5, Q8, FP16)
- Source: Qwen official documentation for llama.cpp
- Source: Qwen3 chat template parsing bug (llama.cpp)
- Source: difference in server/CLI behavior for the reasoning block
- Source: llama.cpp troubleshooting guide (quantization and tool calls)
Why does my Qwen3 GGUF never close the “think” tag?+
Does the --jinja option solve all Qwen3 template problems?+
How do you permanently disable Qwen3's reasoning mode with llama.cpp?+
Why does a recently downloaded Qwen3 GGUF behave differently from an older one?+
Why are my Qwen3 tool calls truncated even with --jinja?+
Does the enable_thinking bug also affect Qwen3, not just Qwen3.5?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.