GLM-5 locally with Ollama: the challenger open-weights
GLM-5 is Zhipu AI’s new open-weight family (Tsinghua University). The 9B and 32B versions offer a permissive license, solid reasoning, decent French, and respectable coding ability. This guide shows how to install GLM-5 locally with Ollama, measures real-world performance on RTX 4070 and 4090, and compares it with Qwen3-32B and DeepSeek V3.2 on the same tasks.
#Why GLM-5?
GLM-5 is not as well known as Qwen3 or DeepSeek, but it is one of the most advanced Chinese open-weight models released in recent months. The 9B targets modest configurations (12 GB of VRAM is enough), while the 32B is aimed at RTX 4090 and Mac Studio. On paper, it claims scores close to GPT-4o-mini on MMLU and HumanEval, at the same size as Qwen3-32B.
In practice, GLM-5 shines in three areas: structured reasoning (a thinking mode similar to Qwen3's), French (noticeably better than Qwen3-9B on complex prompts), and Python code. Where it disappoints: the context window is capped at 128k locally despite a 1M promise, and vision is not integrated (unlike GLM-4V).
#Hardware requirements
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- GLM-5 9B Q4_K_M
- ≈ 5.5 GB VRAM. RTX 3060 12 GB, RTX 4060 8 GB (partial offload), Mac M2 16 GB unified, all RTX 4070+.
- GLM-5 32B Q4_K_M
- ≈ 19 GB VRAM. RTX 4090 24 GB is comfortable, RTX 3090 24 GB is too. On 16 GB (4080), you need to drop to Q3_K_M.
- GLM-5 32B Q5_K_M
- ≈ 22-23 GB VRAM. RTX 4090 fits, but with little headroom. Prefer Q4_K_M if you enable a long context.
- Ollama 0.6.0+
- GLM-5 support was added gradually. A recent version is required for the tokenizer and chat template.
- Disk
- Plan on 6 GB (9B Q4) to 23 GB (32B Q5). More if you keep multiple quantizations in parallel.
#1. Retrieve the model from Hugging Face
Zhipu AI publishes the official weights on Hugging Face under the THUDM organization. Two repositories are relevant to you:
These repositories contain weights in safetensors format (FP16/BF16). For Ollama, you need GGUF. Two options: download a GGUF already converted by the community, or convert it yourself.
#2. GGUF and conversion if necessary
If no official or community GGUF is suitable (exotic variant, specific quantization), conversion goes through llama.cpp. The standard procedure:
#3. Install GLM-5 in Ollama (local)
Once you have the GGUF, Ollama imports it via a Modelfile. Create a file named Modelfile next to the GGUF:
Then create the model in Ollama, verify that it saves, and launch it:
#4. Test quality in French
GLM-5's French is one of its strengths for a model trained primarily on Chinese and English. Three useful prompts for a quick evaluation:
- 01Structured reasoning“A bottle and its cap cost €1.10. The bottle costs €1 more than the cap. How much does the cap cost? Explain your reasoning in detail.” GLM-5 9B should answer €0.05 by showing the equation.
- 02Long-form synthesisPaste a 2,000-word article and ask for a summary in 5 key points in French. The 32B maintains a clear thread, while the 9B shortens the content too much with extended context.
- 03Context-aware Python code“Write a Python function that parses a CSV with a ; delimiter, handles quoted fields, and returns a dict generator. No pandas.” GLM-5 produces clean, idiomatic code, whereas Qwen3-9B tends to import csv without handling quoted fields correctly.
#5. Tokens/sec measured on RTX 4070 and 4090
Measurements taken with Ollama 0.6.x, OLLAMA_FLASH_ATTENTION=1, 8k context, a 500-token prompt, and 300-token generation. Average over 5 runs.
- RTX 4070 12 GB — GLM-5 9B Q4_K_M
- ≈ 55–60 tokens/sec. Comfortable for chat; the model remains 100% in VRAM.
- RTX 4070 12 GB — GLM-5 9B Q5_K_M
- ≈ 48–52 tokens/sec. Slight quality gain, measurable speed loss.
- RTX 4090 24 GB—GLM-5 9B Q4_K_M
- ≈ 110–120 tokens/sec. Vastly overpowered for the 9B.
- RTX 4090 24 GB — GLM-5 32B Q4_K_M
- ≈ 22–26 tokens/sec. Target use case: 32B quality at an interactively usable speed.
- RTX 4090 24 GB — GLM-5 32B Q5_K_M
- ≈ 18–22 tokens/sec. Tight VRAM headroom if context > 16k.
- Mac M4 Pro 48 GB — GLM-5 32B Q4_K_M
- ≈ 12–15 tokens/sec. Acceptable for non-streaming tasks (summarization, analysis).
#GLM-5 vs Qwen3-32B vs DeepSeek V3.2
With the same prompt on the same machine, here’s what we observe in practice. This is a qualitative reading, not a standardized benchmark—use it as a compass, not absolute truth.
- General French
- Qwen3-32B ≥ GLM-5 32B > DeepSeek V3.2 (in fast mode). Qwen3 remains ahead in responsiveness, GLM-5 is very close, and DeepSeek V3.2 hallucinates more in French.
- Structured reasoning
- GLM-5 32B (thinking mode) ≈ Qwen3-32B (thinking). DeepSeek V3.2 is ahead if you let it use its long thinking mode, but at 3–5x the token cost.
- Python code
- Qwen3-Coder > GLM-5 32B > DeepSeek V3.2 ≈ Qwen3-32B. For dedicated coding, stick with Qwen3-Coder. GLM-5 32B is a solid generalist.
- Speed at equal VRAM (24 GB)
- GLM-5 32B Q4 ≈ Qwen3-32B Q4 (~25 tok/s). DeepSeek V3.2 does not fit in 24 GB: it's a 685B MoE, out of contention for this tier.
- Disk footprint
- GLM-5 32B Q4 ≈ 19 GB. Qwen3-32B Q4 ≈ 20 GB. DeepSeek V3.2 ≈ 380 GB Q4—out of the running without a cluster.
#Troubleshooting
- Garbled or entirely Chinese output
- Poorly converted tokenizer or incorrect template. Check the version of llama.cpp used for conversion and make sure the Modelfile includes the <|user|>, <|assistant|>, [gMASK]<sop> tags.
- Stop tokens ignored (infinite response)
- Add PARAMETER stop "<|endoftext|>" and PARAMETER stop "<|user|>" to the Modelfile. Without this, the model may continue with fictional dialogue turns.
- OOM with long context on RTX 4090
- The KV cache explodes beyond 32k. Enable OLLAMA_KV_CACHE_TYPE=q8_0 (about 40% less cache VRAM) or reduce num_ctx.
- Speed < 10 tok/s at 9B on a 4070
- The model spilled into RAM. Check with ollama ps: the PROCESSOR column must show 100% GPU. If CPU appears, lower num_ctx or switch to Q3_K_M.
- Thinking mode too verbose
- GLM-5 can produce long chains of thought. To cut them off, explicitly ask for “a direct answer without intermediate reasoning” in the system prompt.
#GLM-5.2 is out: should you migrate?
Since June 13, 2026, Zhipu has offered GLM-5.2: approximately 753B parameters in MoE, an MIT license, a 1M context, and designed for long-running coding agents. As of today, it is the best-served “frontier” model locally—GGUF quants (unsloth) exist for both Ollama and LM Studio. Plan on a 256 GB Mac or a RTX 4090 machine in 2-bit quantization: if your hardware already handled GLM-5, this guide’s installation steps remain the same; only the model reference changes.
#Go further
GLM-5 deserves some investment to get the most out of it. Three useful directions:
- Choose the right quantization
- Q4_K_M is the default sweet spot, but GLM-5 32B in Q5_K_M fits on 24 GB with a reasonable context and noticeably improves code.
- Customize via Modelfile
- System prompt, response template, sampling parameters: a well-tuned Modelfile saves you from repeating the instructions every session.
- Compare with Qwen3 under real-world conditions
- The Qwen3-32B vs. GLM-5 guide (coming soon) proposes a comparison protocol using 30 representative prompts (French, code, reasoning, summarization).
Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — PC alternative for quantized local LLMs; this isn't the Mac in the tutorial. All AI hardware →
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.