Intermediate 11 minOllama

GLM-5 locally with Ollama: the challenger open-weights

GLM-5 is Zhipu AI’s new open-weight family (Tsinghua University). The 9B and 32B versions offer a permissive license, solid reasoning, decent French, and respectable coding ability. This guide shows how to install GLM-5 locally with Ollama, measures real-world performance on RTX 4070 and 4090, and compares it with Qwen3-32B and DeepSeek V3.2 on the same tasks.

By Mohamed Meguedmi·Update 2026-08-11·Tested on Windows, macOS, and Linux

#Why GLM-5?

GLM-5 is not as well known as Qwen3 or DeepSeek, but it is one of the most advanced Chinese open-weight models released in recent months. The 9B targets modest configurations (12 GB of VRAM is enough), while the 32B is aimed at RTX 4090 and Mac Studio. On paper, it claims scores close to GPT-4o-mini on MMLU and HumanEval, at the same size as Qwen3-32B.

In practice, GLM-5 shines in three areas: structured reasoning (a thinking mode similar to Qwen3's), French (noticeably better than Qwen3-9B on complex prompts), and Python code. Where it disappoints: the context window is capped at 128k locally despite a 1M promise, and vision is not integrated (unlike GLM-4V).

i
GLM-5 vs GLM 5.1
This guide covers GLM-5 (initial release). GLM 5.1 (released later) fixes several regressions in code and improves instruction following. If you’re starting from scratch today, also read the site’s GLM 5.1 guide before deciding.

#Hardware requirements

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
GLM-5 9B Q4_K_M
≈ 5.5 GB VRAM. RTX 3060 12 GB, RTX 4060 8 GB (partial offload), Mac M2 16 GB unified, all RTX 4070+.
GLM-5 32B Q4_K_M
≈ 19 GB VRAM. RTX 4090 24 GB is comfortable, RTX 3090 24 GB is too. On 16 GB (4080), you need to drop to Q3_K_M.
GLM-5 32B Q5_K_M
≈ 22-23 GB VRAM. RTX 4090 fits, but with little headroom. Prefer Q4_K_M if you enable a long context.
Ollama 0.6.0+
GLM-5 support was added gradually. A recent version is required for the tokenizer and chat template.
Disk
Plan on 6 GB (9B Q4) to 23 GB (32B Q5). More if you keep multiple quantizations in parallel.

#1. Retrieve the model from Hugging Face

Zhipu AI publishes the official weights on Hugging Face under the THUDM organization. Two repositories are relevant to you:

Official 9B repository
https://huggingface.co/THUDM/glm-5-9b-chat
Official 32B repository
https://huggingface.co/THUDM/glm-5-32b-chat

These repositories contain weights in safetensors format (FP16/BF16). For Ollama, you need GGUF. Two options: download a GGUF already converted by the community, or convert it yourself.

→
Shortcut: ready-made GGUF
Search Hugging Face for glm-5 GGUF — bartowski, mradermacher, and lmstudio-community publish up-to-date Q4_K_M, Q5_K_M, and Q8_0 conversions. This is the pragmatic option for 95% of cases.

#2. GGUF and conversion if necessary

If no official or community GGUF is suitable (exotic variant, specific quantization), conversion goes through llama.cpp. The standard procedure:

Clone llama.cpp and install the dependencies
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
pip install -r requirements.txt
Convert HF → GGUF FP16
python convert_hf_to_gguf.py /chemin/vers/glm-5-9b-chat \
  --outfile glm-5-9b-chat-f16.gguf \
  --outtype f16
Quantize in Q4_K_M
./build/bin/llama-quantize \
  glm-5-9b-chat-f16.gguf \
  glm-5-9b-chat-Q4_K_M.gguf \
  Q4_K_M
!
GLM Tokenizer
GLM uses a derived SentencePiece tokenizer. If convert_hf_to_gguf.py complains about missing special tokens, update llama.cpp (GLM-5 support was added in a recent PR). With a version that is too old, conversion succeeds but inference produces gibberish.

#3. Install GLM-5 in Ollama (local)

Once you have the GGUF, Ollama imports it via a Modelfile. Create a file named Modelfile next to the GGUF:

Modelfile for GLM-5 9B
FROM ./glm-5-9b-chat-Q4_K_M.gguf

PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 32768
PARAMETER stop "<|user|>"
PARAMETER stop "<|endoftext|>"

TEMPLATE """[gMASK]<sop>{{ if .System }}<|system|>
{{ .System }}{{ end }}<|user|>
{{ .Prompt }}<|assistant|>
{{ .Response }}"""

SYSTEM """Tu es un assistant utile, précis et concis. Tu réponds en français sauf demande contraire."""

Then create the model in Ollama, verify that it saves, and launch it:

Creation and launch
ollama create glm5-9b -f Modelfile
ollama list
ollama run glm5-9b
→
When native support arrives
If Ollama publishes GLM-5 in its official registry (ollama run glm5 without a Modelfile), prefer it: the chat template will be maintained upstream, and you’ll avoid stop-token errors.

#4. Test quality in French

GLM-5's French is one of its strengths for a model trained primarily on Chinese and English. Three useful prompts for a quick evaluation:

  1. 01
    Structured reasoning
    “A bottle and its cap cost €1.10. The bottle costs €1 more than the cap. How much does the cap cost? Explain your reasoning in detail.” GLM-5 9B should answer €0.05 by showing the equation.
  2. 02
    Long-form synthesis
    Paste a 2,000-word article and ask for a summary in 5 key points in French. The 32B maintains a clear thread, while the 9B shortens the content too much with extended context.
  3. 03
    Context-aware Python code
    “Write a Python function that parses a CSV with a ; delimiter, handles quoted fields, and returns a dict generator. No pandas.” GLM-5 produces clean, idiomatic code, whereas Qwen3-9B tends to import csv without handling quoted fields correctly.
i
Thinking mode
GLM-5 has a chain-of-thought mode that can be enabled via a system prefix (see the model card on HF for the exact format). When enabled, reasoning quality improves noticeably—but at the cost of about 30-40% more tokens.

#5. Tokens/sec measured on RTX 4070 and 4090

Measurements taken with Ollama 0.6.x, OLLAMA_FLASH_ATTENTION=1, 8k context, a 500-token prompt, and 300-token generation. Average over 5 runs.

RTX 4070 12 GB — GLM-5 9B Q4_K_M
≈ 55–60 tokens/sec. Comfortable for chat; the model remains 100% in VRAM.
RTX 4070 12 GB — GLM-5 9B Q5_K_M
≈ 48–52 tokens/sec. Slight quality gain, measurable speed loss.
RTX 4090 24 GB—GLM-5 9B Q4_K_M
≈ 110–120 tokens/sec. Vastly overpowered for the 9B.
RTX 4090 24 GB — GLM-5 32B Q4_K_M
≈ 22–26 tokens/sec. Target use case: 32B quality at an interactively usable speed.
RTX 4090 24 GB — GLM-5 32B Q5_K_M
≈ 18–22 tokens/sec. Tight VRAM headroom if context > 16k.
Mac M4 Pro 48 GB — GLM-5 32B Q4_K_M
≈ 12–15 tokens/sec. Acceptable for non-streaming tasks (summarization, analysis).
→
Enable Flash Attention
Setting OLLAMA_FLASH_ATTENTION=1 in the environment delivers a 10-15% gain on GLM-5 32B and stabilizes VRAM during long contexts. Combined with OLLAMA_KV_CACHE_TYPE=q8_0, it frees up an additional 2-3 GB with no visible degradation.

#GLM-5 vs Qwen3-32B vs DeepSeek V3.2

With the same prompt on the same machine, here’s what we observe in practice. This is a qualitative reading, not a standardized benchmark—use it as a compass, not absolute truth.

General French
Qwen3-32B ≥ GLM-5 32B > DeepSeek V3.2 (in fast mode). Qwen3 remains ahead in responsiveness, GLM-5 is very close, and DeepSeek V3.2 hallucinates more in French.
Structured reasoning
GLM-5 32B (thinking mode) ≈ Qwen3-32B (thinking). DeepSeek V3.2 is ahead if you let it use its long thinking mode, but at 3–5x the token cost.
Python code
Qwen3-Coder > GLM-5 32B > DeepSeek V3.2 ≈ Qwen3-32B. For dedicated coding, stick with Qwen3-Coder. GLM-5 32B is a solid generalist.
Speed at equal VRAM (24 GB)
GLM-5 32B Q4 ≈ Qwen3-32B Q4 (~25 tok/s). DeepSeek V3.2 does not fit in 24 GB: it's a 685B MoE, out of contention for this tier.
Disk footprint
GLM-5 32B Q4 ≈ 19 GB. Qwen3-32B Q4 ≈ 20 GB. DeepSeek V3.2 ≈ 380 GB Q4—out of the running without a cluster.
i
Practical verdict
If you have 24 GB of VRAM and want a general-purpose 32B model, the trade-off is between GLM-5 and Qwen3. GLM-5 is more interesting for structured reasoning and for venturing somewhat off the beaten path; Qwen3-32B remains the safe choice. DeepSeek V3.2 remains an API or large-server model—not a question at 24 GB.

#Troubleshooting

Garbled or entirely Chinese output
Poorly converted tokenizer or incorrect template. Check the version of llama.cpp used for conversion and make sure the Modelfile includes the <|user|>, <|assistant|>, [gMASK]<sop> tags.
Stop tokens ignored (infinite response)
Add PARAMETER stop "<|endoftext|>" and PARAMETER stop "<|user|>" to the Modelfile. Without this, the model may continue with fictional dialogue turns.
OOM with long context on RTX 4090
The KV cache explodes beyond 32k. Enable OLLAMA_KV_CACHE_TYPE=q8_0 (about 40% less cache VRAM) or reduce num_ctx.
Speed < 10 tok/s at 9B on a 4070
The model spilled into RAM. Check with ollama ps: the PROCESSOR column must show 100% GPU. If CPU appears, lower num_ctx or switch to Q3_K_M.
Thinking mode too verbose
GLM-5 can produce long chains of thought. To cut them off, explicitly ask for “a direct answer without intermediate reasoning” in the system prompt.

#GLM-5.2 is out: should you migrate?

Since June 13, 2026, Zhipu has offered GLM-5.2: approximately 753B parameters in MoE, an MIT license, a 1M context, and designed for long-running coding agents. As of today, it is the best-served “frontier” model locally—GGUF quants (unsloth) exist for both Ollama and LM Studio. Plan on a 256 GB Mac or a RTX 4090 machine in 2-bit quantization: if your hardware already handled GLM-5, this guide’s installation steps remain the same; only the model reference changes.

i
Dedicated guide coming soon
A complete GLM-5.2 guide (installation Ollama + LM Studio, realistic configurations, and honest expectations for speeds) is in preparation and will be published on the site this week.

#Go further

GLM-5 deserves some investment to get the most out of it. Three useful directions:

Choose the right quantization
Q4_K_M is the default sweet spot, but GLM-5 32B in Q5_K_M fits on 24 GB with a reasonable context and noticeably improves code.
Customize via Modelfile
System prompt, response template, sampling parameters: a well-tuned Modelfile saves you from repeating the instructions every session.
Compare with Qwen3 under real-world conditions
The Qwen3-32B vs. GLM-5 guide (coming soon) proposes a comparison protocol using 30 representative prompts (French, code, reasoning, summarization).

Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — PC alternative for quantized local LLMs; this isn't the Mac in the tutorial. All AI hardware →

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.