GLM 5.1 locally: the open-weight alternative to know
GLM 5.1 is the latest iteration of Zhipu AI's open-weight model family, from one of China's most active labs in self-hosting. The model is often overshadowed by Qwen and Llama in French-language discussions, unfairly so: for some use cases — coding, structured reasoning, and tool calling — it holds its own. This guide shows you how to install GLM 5.1 locally with Ollama or llama.cpp, how much VRAM it uses at each quantization level, how it performs in French and coding, and where it stands against Qwen 3.5 and Gemma 4.
#Why run GLM 5.1 locally
Three concrete reasons favor trying GLM 5.1 over an equivalent Qwen 3.5 or Gemma 4. First, its permissive license: the open-weight version (available in 9B and 32B) permits commercial use, which is far from obvious in the open-source ecosystem. Second, native tool calling: GLM 5.1 was trained from the outset to call tools, making it particularly well suited to local agents without prompt hacks.
Finally, the quality-to-size ratio: the 9B reaches the level of a Gemma 4 12B on several code and reasoning benchmarks, making it an interesting choice if you're limited to 12 GB of VRAM. The 32B aims higher and competes with 24–35B models such as Qwen 3.8 27B or Mistral Small 24B.
#What defines GLM 5.1
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
GLM 5.1 retains the dense transformer architecture of previous versions—no Mixture-of-Experts here, unlike Qwen 3.6 35B-A3B or DeepSeek V4. This choice simplifies inference and ensures that throughput doesn't collapse on questions that engage unusual experts.
- Sizes
- 9B and 32B in open-weight form. A 100B+ version exists, but it is available only through the Zhipu API and cannot be downloaded.
- Context
- 128k tokens at both sizes, extended to 1M via YaRN scaling on the 32B (with degradation beyond 200k).
- Tokenizer
- Optimized Chinese/English bilingual tokenizer, with decent French coverage. Expect ~5% more tokens than a Mistral for an equivalent French text.
- Tool calling
- Native format compatible with the OpenAI tools schema. No prompt engineering is needed to trigger it; the model knows when to call a function.
- License
- GLM License (an Apache 2.0 variant with an attribution clause). Commercial use permitted; redistribution under the same terms.
- Multimodal
- Text-only on the 9B and 32B open-weight models. The vision version (GLM-V) is a separate model that must be pulled separately.
#Prerequisites
- Ollama ≥ 0.5
- For version Ollama. Older builds do not know the GLM 5.1 chat template.
- Recent llama.cpp
- 2026 build or later. Support for the GLM architecture was stabilized in late 2025.
- Disk space
- 6 GB for the 9B Q4, 19 GB for the 32B Q4. Double that for Q8, quadruple it for FP16.
- Minimum VRAM
- 8 GB for the 9B Q4, 24 GB for the 32B Q4. Below that, it spills over to the CPU and throughput drops to 3–5 tok/s.
- Up-to-date GPU driver
- CUDA 12.x for NVIDIA, ROCm 6.x for AMD, native Metal for Apple Silicon.
#1. Installation with Ollama
Ollama is the shortest path. The model is served under the name glm5.1 in the official library, with the usual tags for size and quantization.
For the 32B, the mechanics are the same, but allow several minutes for the download and at least 24 GB of VRAM or unified memory.
Once pulled, Ollama's OpenAI-compatible endpoint listens by default on http://localhost:11434, and glm5.1 responds like any model served by Ollama.
#2. Installation with llama.cpp
With llama.cpp, you retrieve a GGUF from Hugging Face and run it in the CLI or through the HTTP server. This is the preferred approach if you want fine-grained control over the KV cache, the draft model for speculation, or if you are running on multi-GPU.
- -c 16384
- Context window. 16k is a good default for French/code; increase it to 32k or more if you have the VRAM.
- -ngl 99
- All on the GPU. Use a lower value to offload layers to the CPU if you run out of room.
- -fa
- Flash Attention. Essential beyond 8k to prevent KV memory from exploding.
- --host 0.0.0.0
- Exposes the server on the network. Set it to 127.0.0.1 if you want to restrict it to the local machine.
#3. VRAM by quantization
The figures below count the weights plus a KV cache for 8k context. To increase this to 32k or 128k, add 2 to 8 GB depending on the size.
- GLM 5.1 9B — Q4_K_M
- ≈ 6 GB VRAM. Fits on GTX 1660 6 GB with tight margins, comfortable on RTX 3050/3060/4060 8 GB.
- GLM 5.1 9B — Q5_K_M
- ≈ 7 GB VRAM. Relevant from 8 GB with some headroom.
- GLM 5.1 9B — Q8_0
- ≈ 10 GB VRAM. Target RTX 3060 12 GB, 4070 12 GB, or Mac 16 GB+.
- GLM 5.1 32B — Q4_K_M
- ≈ 19 GB VRAM. Sweet spot: RTX 3090, 4090, 5090, RX 7900 XTX, Mac Studio.
- GLM 5.1 32B — Q5_K_M
- ≈ 23 GB VRAM. At the upper limit of 24 GB; target a Mac with 48 GB+ or multi-GPU to have room to breathe.
- GLM 5.1 32B — Q8_0
- ≈ 34 GB VRAM. Reserved for the RTX 5090 32 GB, Mac Studio Ultra systems, or multi-GPU setups (2× 3090, 2× 4090).
#4. Quality in French and code
GLM 5.1 is trained mainly on Chinese and English, but French works reasonably well—not at the level of a native Mistral, but above a Gemma 4 12B in the same size class. Typical errors involve idiomatic phrasing and a few uncommon agreements; nothing that interferes with general assistant use or RAG.
- French — 9B
- Strong for summarization, translation, and RAG on French documents. A few occasional Anglicisms. Prefer a low temperature (0.2-0.4) to minimize drift.
- French — 32B
- Very polished. Neck and neck with Mistral Small 24B for writing quality, slightly below the largest proprietary models.
- Code — 9B
- Excellent for its size. High-quality Python, JS, Go, and Rust generation. Beats a general-purpose Granite 4.2 8B and holds its own against Qwen 3.5 9B on code.
- Code — 32B
- Very strong performance. On HumanEval, MBPP, and MultiPL-E, it ranks among the leading open-weight models. Good at multi-file refactoring thanks to its 128k context.
- Reasoning
- Explicit chain-of-thought when requested, without a special <think> token like DeepSeek R1. The 32B performs well on AIME, GPQA, and MATH.
- Tool calling
- Native format compatible with OpenAI tools. The model decides on its own to call a function and formats the JSON arguments cleanly.
#5. GLM 5.1 vs Qwen 3.5 and Gemma 4
Three open-weight families occupy the same size class in 2026: GLM 5.1, Qwen (3.5 9B / 3.8 27B), and Gemma 4 (12B / 26B-A4B). The honest comparison:
- Raw speed (tok/s)
- At the same size and VRAM, GLM 5.1 falls in the same range as Qwen 3.5 dense and Gemma 4. No meaningful gap—the throughput depends mainly on the backend (Ollama vs llama.cpp vs vLLM) and the GPU.
- French-language quality
- Mistral Small 24B native > Qwen 3.5 ≈ GLM 5.1 > Gemma 4 at equivalent size. If French is critical, Mistral remains ahead; GLM 5.1 is a strong second open-model choice.
- Code
- Qwen3-Coder 30B / Devstral 24B > GLM 5.1 ≈ Qwen 3.8 general-purpose > Gemma 4 at the same size. GLM 5.1 32B is competitive with Qwen 3.8 27B, slightly behind dedicated code specialists.
- Tool calling
- GLM 5.1 ≈ Qwen 3.5 > Gemma 4. The first two were explicitly trained for tools, while Gemma requires a bit more prompt engineering.
- Long context
- Qwen 3.5 with native 256k (Qwen 3.8 27B: 262k) > GLM 5.1 128k (1M via YaRN) > Gemma 4 128k. For whole-codebase analysis, Qwen remains ahead.
- License
- GLM (Apache-like) ≈ Qwen (Apache 2.0 for most variants) ≈ Gemma 4 (Apache 2.0 since April 2026). Three straightforward licenses for enterprise use, with Gemma having dropped its custom clause.
- Ecosystem
- Qwen > Gemma 4 > GLM 5.1. Qwen now has the largest body of community fine-tunes, Gemma 4 is close behind, while GLM 5.1 remains more niche—fewer variants, fewer French tutorials.
#Tips and pitfalls
- Short system prompt
- GLM 5.1 follows system prompts well but gets a bit saturated if you give it 800 words. Aim for 200–400 words, clear and structured.
- Low temperature for FR
- 0.2–0.4 to minimize anglicisms and agreement oddities. Above 0.7, the quality of the French writing starts to fluctuate.
- Stop tokens
- The model uses <|endoftext|> and <|user|> as delimiters. If you write your own client, add these two strings to the stop tokens to prevent it from continuing to talk on its own.
- Streaming via Ollama
- The /api/chat endpoint and /v1/chat/completions support stream=true. There is no particular quirk with GLM 5.1, unlike some MoE models.
- LoRA fine-tuning
- GLM 5.1 9B is viable with LoRA on a RTX 4090 24 GB using Unsloth. The 32B requires multi-GPU or a Mac with generous unified memory.
- Ollama Modelfile
- To lock in a system prompt and parameters, create a Modelfile based on glm5.1:9b, then ollama create monassistant -f Modelfile. The derived model appears in ollama list like any other.
#Go further
GLM 5.1 is running, and you have your first results. A few ways to push it further:
- Choose your quantization (Q4, Q5, Q8, FP16)
- For precisely deciding the quality/VRAM tradeoff for your use cases.
- Qwen3.6 35B-A3B locally: testing and VRAM requirements
- The direct competitor on the Qwen side, using an MoE architecture. Useful comparison if you’re undecided.
- Open WebUI with Ollama: complete guide
- A ChatGPT-like interface on top of GLM 5.1: conversations, RAG, and multiple users.
- Build a local AI agent in Python with LangChain and Ollama
- To use GLM 5.1's native tool calling and have it call your own functions.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.