Advanced 11 minQuantization

Qwen3 GGUF: fix the tokenizer and chat template

Direct response

Tokenizer and chat-template errors in Qwen3 GGUFs appear as three symptoms: a reasoning tag that never closes, generation that never stops, or irrelevant answers. The cause is almost always an incorrectly applied chat template. Add --jinja to use the template embedded in the GGUF, verify the file's provenance, and as a last resort pass a custom template.

A GGUF Qwen3 that “answers nonsense” is almost never a model quality issue: it’s a prompt formatting issue before the prompt reaches the model. This guide covers only Qwen3 GGUF tokenizer and chat template errors: recognizing the symptom, checking the file’s provenance, and fixing or replacing the template.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#The three symptoms to recognize

Three signals point to a tokenizer or chat-template problem rather than a model problem: a thinking tag (usually “think”) that opens and never closes in the response, generation that continues indefinitely without stopping at the logical end of the response, or a model that responds off-topic as if it did not understand that it was being asked a question. In all three cases, the model itself is not at fault: the structure of the text it receives as input is malformed.

i
Why this happens specifically with Qwen3
Qwen3’s chat template handles a more complex structure than most earlier models: switching between thinking and direct modes, tool calls, and Jinja syntax with constructs such as list slicing that were not all supported by early versions of llama.cpp’s template engine.

#The root cause: the chat template

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A language model never receives your messages directly: a chat template formats them (turn tags, system prompt, start/end markers) before transforming them into tokens. If this template is missing, incorrectly chosen, or misinterpreted by the inference engine, the model receives text that does not resemble what it was trained on and produces degraded outputs, even if the model weights and tokenizer themselves are correct.

#The --jinja option: the first thing to check

The official Qwen documentation explicitly recommends adding --jinja when launching a Qwen3 GGUF with llama.cpp: this option tells it to use the chat template embedded in the GGUF file, presented as the preferred method over a generic template selected by the engine by default.

Reference command documented by Qwen
./llama-cli -hf Qwen/Qwen3-8B-GGUF:Q8_0 --jinja --color -ngl 99 -fa -sm row --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0 -c 40960 -n 32768 --no-context-shift

If your launch command does not contain --jinja, this is the first fix to try before considering any other hypothesis. Many scripts and interfaces built before this option became widespread still omit it, which explains a large share of reports of “broken responses” from otherwise valid Qwen3 GGUFs.

#A known parsing bug, now fixed

A specific error affected llama.cpp's Qwen3 chat template: the template engine failed to parse a Jinja list-slicing syntax (messages[::-1]) used to traverse the conversation history backward in the tool-calling logic. The reported error was a parsing failure pointing precisely to that line of the template.

!
Error message to recognize
“Expected value expression at row 18, column 30” followed by a line containing messages[::-1] in the template: that's the exact signature of this parsing bug. It was fixed by a later pull request in the llama.cpp repository; if you still encounter it, the first thing to check is your llama.cpp binary version, which is probably older than the fix.
  1. 01
    Identify the llama.cpp version in use
    An older version may not include the parsing fix for the slicing syntax used by the Qwen3 template.
  2. 02
    Update to a recent version
    Recompiling or redownloading an up-to-date llama.cpp binary resolves this specific case without modifying the GGUF itself.
  3. 03
    If updating isn't possible
    Pass a simplified custom template via --chat-template-file, avoiding the problematic Jinja construction.

#llama-server and llama-cli do not behave the same way

A behavior reported in the llama.cpp repository: enabling --jinja with llama-server can make the reasoning block—the content between the thinking tags—disappear from the response, while the same block remains visible when using llama-cli with the same option and model. If your integration depends on reasoning content being present in the output, for observability or debugging, this isn't a tokenizer issue but a processing difference between the two binaries—check which one you're using before investigating further.

This distinction has a practical consequence for anyone building an integration on top of llama.cpp rather than a simple interactive session: a test pipeline that validates the output format with llama-cli and is then deployed in production behind llama-server may show a silent regression on this specific point even though no application-side parameter has changed. Explicitly documenting which binary serves production, and testing with that exact binary rather than the one used in local development, avoids this trap.

#Disable forced reasoning mode

Qwen3 provides a mechanism for switching between thinking and direct modes at the chat template level. However, the official documentation Qwen indicates that this forced deactivation mechanism (hard switch) is not natively exposed in llama.cpp: passing enable_thinking as false through the command-line options may be ignored depending on the version, as shown by several recent reports involving Qwen3.5 variants.

The workaround documented by Qwen is to provide a custom template via --chat-template-file, with enable_thinking explicitly set to false at the template level itself rather than passed as a parameter at request time. This is more reliable than a runtime parameter that depends on the exact support in your version of llama.cpp.

One point needs clarification to avoid overgeneralizing: report #20182 ("enable_thinking param cannot turn off thinking") specifically concerns Qwen3.5-9B in build 8215, remains labeled "bug-unconfirmed" in the llama.cpp repository, and was closed without resolution ("not planned"). There is no evidence that the same behavior affects an original Qwen3 GGUF (as opposed to Qwen3.5): if you encounter this symptom on a standard Qwen3, treat it as a case to isolate and report separately rather than as automatic confirmation of this ticket.

#Verify the provenance of a third-party GGUF

Some tokenizer issues with Qwen3 GGUF files do not come from llama.cpp but from the GGUF file itself: a conversion made with an old version of the conversion tools, or a file whose tokenizer was exported incorrectly, produces similar symptoms (missing end of generation, special tokens not recognized correctly). Before looking for a bug in the inference engine, comparing your GGUF file's size and publication date with a recognized repository (Qwen official, or documented requantizations) can rule out this hypothesis.

A simple way to quickly distinguish a file problem from a configuration problem: if the same GGUF works correctly on another machine, or with another version of llama.cpp, the file itself is probably not at fault. Conversely, if multiple re-downloads of the same repository on the same machine consistently reproduce the symptom, the most likely explanation shifts toward the local configuration (binary version, launch options) rather than the file.

→
Useful reflex
If a recently downloaded GGUF shows tokenizer errors that an older GGUF of the same model did not, download the file again from its original source before suspecting a llama.cpp bug: a corrupted download or an incomplete conversion are common causes that are easy to rule out.

#Quantization that is too low: malformed tool calls

A final symptom, distinct from the first three, specifically affects using Qwen3 as an agent with tool calls: instead of a text-formatting problem, the tool call itself arrives truncated, with empty or malformed arguments (invalid JSON). A community troubleshooting guide dedicated to llama.cpp documents that tool-call structure is sensitive to the quantization level: sub-4-bit quantizations (Q3, Q2, IQ) produce malformed tool calls even when the chat template is correctly applied with --jinja.

!
Check before blaming the template
If your tool calls work correctly with a GGUF Q5_K_M or Q6_K but break on the same model in Q3 or Q2, the cause is not the chat template but the quantization itself: reduced weight precision affects the generation of strict structures (JSON, tags) more than overall text quality. The documented fix is to use a higher quantization level, not to modify the template.

This issue is easy to miss because, at first glance, it looks like a classic tokenizer problem: a truncated response immediately suggests an improperly closed chat template. The practical distinction lies in the context where the symptom appears: a template problem affects every response, including plain text without a tool call, whereas a quantization problem affecting tool calls generally leaves free-form text responses untouched and appears only in the strict JSON structure expected by the tool-calling protocol.

#Quick troubleshooting table

Observed symptom, most likely cause, first correction to try
SymptomMost likely causeCorrection first
Never-closed « think » tagChat template not appliedAdd --jinja when launching
Generation that never stopsTemplate misinterpreted or missingCheck --jinja; otherwise update llama.cpp
“Expected value expression” error on launchBug in Jinja slicing parsing (fixed by PR #13573)Update to a recent version of llama.cpp
Reasoning block absent with llama-server but visible with llama-cliDocumented processing difference between the two binariesTest with llama-cli to confirm, then follow ticket #14894
enable_thinking=false ignoredHard switch not natively exposed in llama.cppSet enable_thinking=false in a template via --chat-template-file
Truncated tool calls or invalid JSONQuantization too low (Q3, Q2, IQ)Move up to at least Q4_K_M, ideally Q5_K_M or Q6_K
Newer GGUF more buggy than an older one for the same modelPoorly converted file or corrupted downloadDownload again from the original source (official Qwen or recognized repository)
Frequently asked questions
Why does my Qwen3 GGUF never close the “think” tag?+
This is the most common symptom of an incorrectly applied chat template. First check that your command includes --jinja to use the template embedded in the GGUF rather than a generic default template, which is missing from many scripts and interfaces built before this option became widespread.
Does the --jinja option solve all Qwen3 template problems?+
It handles most cases, but not all: a parsing bug involving list-slicing syntax affected some versions of llama.cpp (fixed since), behavior sometimes differs between llama-server and llama-cli when displaying the reasoning block, and overly low quantization can break tool calls despite a correct template.
How do you permanently disable Qwen3's reasoning mode with llama.cpp?+
The enable_thinking parameter passed on the command line may be ignored depending on the version. The method documented by Qwen is to provide a custom template via --chat-template-file, with enable_thinking set to false directly in the template, rather than passed as a request parameter whose support depends on your exact build.
Why does a recently downloaded Qwen3 GGUF behave differently from an older one?+
First, verify the file’s provenance and integrity (redownload it from the official source Qwen or a recognized repository) before suspecting a llama.cpp bug: an incomplete GGUF conversion or a corrupted download can produce tokenizer symptoms very similar to a template bug, but redownloading the file fixes them.
Why are my Qwen3 tool calls truncated even with --jinja?+
A community troubleshooting guide documents that tool-call structure is sensitive to quantization: sub-4-bit quantizations (Q3, Q2, IQ) produce malformed JSON even with the correct template. Moving up to Q4_K_M or a higher quantization (Q5_K_M, Q6_K) is the first correction to test.
Does the enable_thinking bug also affect Qwen3, not just Qwen3.5?+
The best-documented reports (issues #20182 and #20409) explicitly concern Qwen3.5 variants and remain marked “bug-unconfirmed,” closed without resolution. Nothing confirms the same behavior on an original Qwen3 GGUF: verify it case by case rather than assuming it is identical.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.