Intermediate 10 minGLM

GLM 5.1 locally: the open-weight alternative to know

GLM 5.1 is the latest iteration of Zhipu AI's open-weight model family, from one of China's most active labs in self-hosting. The model is often overshadowed by Qwen and Llama in French-language discussions, unfairly so: for some use cases — coding, structured reasoning, and tool calling — it holds its own. This guide shows you how to install GLM 5.1 locally with Ollama or llama.cpp, how much VRAM it uses at each quantization level, how it performs in French and coding, and where it stands against Qwen 3.5 and Gemma 4.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run GLM 5.1 locally

Three concrete reasons favor trying GLM 5.1 over an equivalent Qwen 3.5 or Gemma 4. First, its permissive license: the open-weight version (available in 9B and 32B) permits commercial use, which is far from obvious in the open-source ecosystem. Second, native tool calling: GLM 5.1 was trained from the outset to call tools, making it particularly well suited to local agents without prompt hacks.

Finally, the quality-to-size ratio: the 9B reaches the level of a Gemma 4 12B on several code and reasoning benchmarks, making it an interesting choice if you're limited to 12 GB of VRAM. The 32B aims higher and competes with 24–35B models such as Qwen 3.8 27B or Mistral Small 24B.

i
In two words
GLM 5.1 = a serious Chinese open-weight model, underrated in the French-speaking world, particularly strong at coding and tool calling. Ollama exposes it with one command; llama.cpp if you want more control.

#What defines GLM 5.1

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

GLM 5.1 retains the dense transformer architecture of previous versions—no Mixture-of-Experts here, unlike Qwen 3.6 35B-A3B or DeepSeek V4. This choice simplifies inference and ensures that throughput doesn't collapse on questions that engage unusual experts.

Sizes
9B and 32B in open-weight form. A 100B+ version exists, but it is available only through the Zhipu API and cannot be downloaded.
Context
128k tokens at both sizes, extended to 1M via YaRN scaling on the 32B (with degradation beyond 200k).
Tokenizer
Optimized Chinese/English bilingual tokenizer, with decent French coverage. Expect ~5% more tokens than a Mistral for an equivalent French text.
Tool calling
Native format compatible with the OpenAI tools schema. No prompt engineering is needed to trigger it; the model knows when to call a function.
License
GLM License (an Apache 2.0 variant with an attribution clause). Commercial use permitted; redistribution under the same terms.
Multimodal
Text-only on the 9B and 32B open-weight models. The vision version (GLM-V) is a separate model that must be pulled separately.
→
Why not MoE here
A dense 32B is simpler to deploy than a 35B-A3B MoE: no routing-related VRAM surprises and no major throughput variance depending on the question. If you want stable, predictable speed, GLM 5.1 32B is more comfortable than an equivalent MoE.

#Prerequisites

Ollama ≥ 0.5
For version Ollama. Older builds do not know the GLM 5.1 chat template.
Recent llama.cpp
2026 build or later. Support for the GLM architecture was stabilized in late 2025.
Disk space
6 GB for the 9B Q4, 19 GB for the 32B Q4. Double that for Q8, quadruple it for FP16.
Minimum VRAM
8 GB for the 9B Q4, 24 GB for the 32B Q4. Below that, it spills over to the CPU and throughput drops to 3–5 tok/s.
Up-to-date GPU driver
CUDA 12.x for NVIDIA, ROCm 6.x for AMD, native Metal for Apple Silicon.

#1. Installation with Ollama

Ollama is the shortest path. The model is served under the name glm5.1 in the official library, with the usual tags for size and quantization.

Direct pull and launch
ollama run glm5.1:9b

For the 32B, the mechanics are the same, but allow several minutes for the download and at least 24 GB of VRAM or unified memory.

Useful tags
ollama pull glm5.1:9b           # alias de 9b-instruct-q4_K_M
ollama pull glm5.1:9b-q8_0      # qualité supérieure, +60% VRAM
ollama pull glm5.1:32b          # 32B Q4 par défaut
ollama pull glm5.1:32b-q5_K_M   # le sweet spot 32B sur 24 Go

Once pulled, Ollama's OpenAI-compatible endpoint listens by default on http://localhost:11434, and glm5.1 responds like any model served by Ollama.

i
Verify that the model fits in VRAM
After the first prompt, run ollama ps in another terminal. The PROCESSOR column should show 100% GPU. If you see 70% GPU / 30% CPU, you’ve exceeded the limit—switch to more aggressive quantization or move down one size.

#2. Installation with llama.cpp

With llama.cpp, you retrieve a GGUF from Hugging Face and run it in the CLI or through the HTTP server. This is the preferred approach if you want fine-grained control over the KV cache, the draft model for speculation, or if you are running on multi-GPU.

GGUF download
# 9B en Q4_K_M (recommandé)
wget https://huggingface.co/zhipuai/glm-5.1-9b-chat-gguf/resolve/main/glm-5.1-9b-chat-Q4_K_M.gguf

# 32B en Q4_K_M
wget https://huggingface.co/zhipuai/glm-5.1-32b-chat-gguf/resolve/main/glm-5.1-32b-chat-Q4_K_M.gguf
Launch in server mode
./llama-server \
  -m glm-5.1-9b-chat-Q4_K_M.gguf \
  -c 16384 \
  -ngl 99 \
  -fa \
  --host 0.0.0.0 --port 8080
-c 16384
Context window. 16k is a good default for French/code; increase it to 32k or more if you have the VRAM.
-ngl 99
All on the GPU. Use a lower value to offload layers to the CPU if you run out of room.
-fa
Flash Attention. Essential beyond 8k to prevent KV memory from exploding.
--host 0.0.0.0
Exposes the server on the network. Set it to 127.0.0.1 if you want to restrict it to the local machine.
→
KV-cache quantization
On GLM 5.1, --cache-type-k q8_0 --cache-type-v q8_0 cuts KV memory by 2 with almost no quality loss. Essential if you're aiming for 64k+ of context on a 32B model.

#3. VRAM by quantization

The figures below count the weights plus a KV cache for 8k context. To increase this to 32k or 128k, add 2 to 8 GB depending on the size.

GLM 5.1 9B — Q4_K_M
≈ 6 GB VRAM. Fits on GTX 1660 6 GB with tight margins, comfortable on RTX 3050/3060/4060 8 GB.
GLM 5.1 9B — Q5_K_M
≈ 7 GB VRAM. Relevant from 8 GB with some headroom.
GLM 5.1 9B — Q8_0
≈ 10 GB VRAM. Target RTX 3060 12 GB, 4070 12 GB, or Mac 16 GB+.
GLM 5.1 32B — Q4_K_M
≈ 19 GB VRAM. Sweet spot: RTX 3090, 4090, 5090, RX 7900 XTX, Mac Studio.
GLM 5.1 32B — Q5_K_M
≈ 23 GB VRAM. At the upper limit of 24 GB; target a Mac with 48 GB+ or multi-GPU to have room to breathe.
GLM 5.1 32B — Q8_0
≈ 34 GB VRAM. Reserved for the RTX 5090 32 GB, Mac Studio Ultra systems, or multi-GPU setups (2× 3090, 2× 4090).
!
The 128k context trap
Enabling the maximum context adds 5 to 10 GB to the KV cache, depending on the size. On the 32B Q5 with 24 GB, you can manage 8k but exceed the limit at 32k. Increase the context in stages and monitor ollama ps or nvidia-smi.

#4. Quality in French and code

GLM 5.1 is trained mainly on Chinese and English, but French works reasonably well—not at the level of a native Mistral, but above a Gemma 4 12B in the same size class. Typical errors involve idiomatic phrasing and a few uncommon agreements; nothing that interferes with general assistant use or RAG.

French — 9B
Strong for summarization, translation, and RAG on French documents. A few occasional Anglicisms. Prefer a low temperature (0.2-0.4) to minimize drift.
French — 32B
Very polished. Neck and neck with Mistral Small 24B for writing quality, slightly below the largest proprietary models.
Code — 9B
Excellent for its size. High-quality Python, JS, Go, and Rust generation. Beats a general-purpose Granite 4.2 8B and holds its own against Qwen 3.5 9B on code.
Code — 32B
Very strong performance. On HumanEval, MBPP, and MultiPL-E, it ranks among the leading open-weight models. Good at multi-file refactoring thanks to its 128k context.
Reasoning
Explicit chain-of-thought when requested, without a special <think> token like DeepSeek R1. The 32B performs well on AIME, GPQA, and MATH.
Tool calling
Native format compatible with OpenAI tools. The model decides on its own to call a function and formats the JSON arguments cleanly.
→
For everyday coding
If your primary goal is coding assistance, compare GLM 5.1 9B with Qwen 3.5 9B on your own tasks. Both are excellent, but one may fit your language and style better. No public benchmark can replace 30 minutes of testing on your repo.

#5. GLM 5.1 vs Qwen 3.5 and Gemma 4

Three open-weight families occupy the same size class in 2026: GLM 5.1, Qwen (3.5 9B / 3.8 27B), and Gemma 4 (12B / 26B-A4B). The honest comparison:

Raw speed (tok/s)
At the same size and VRAM, GLM 5.1 falls in the same range as Qwen 3.5 dense and Gemma 4. No meaningful gap—the throughput depends mainly on the backend (Ollama vs llama.cpp vs vLLM) and the GPU.
French-language quality
Mistral Small 24B native > Qwen 3.5 ≈ GLM 5.1 > Gemma 4 at equivalent size. If French is critical, Mistral remains ahead; GLM 5.1 is a strong second open-model choice.
Code
Qwen3-Coder 30B / Devstral 24B > GLM 5.1 ≈ Qwen 3.8 general-purpose > Gemma 4 at the same size. GLM 5.1 32B is competitive with Qwen 3.8 27B, slightly behind dedicated code specialists.
Tool calling
GLM 5.1 ≈ Qwen 3.5 > Gemma 4. The first two were explicitly trained for tools, while Gemma requires a bit more prompt engineering.
Long context
Qwen 3.5 with native 256k (Qwen 3.8 27B: 262k) > GLM 5.1 128k (1M via YaRN) > Gemma 4 128k. For whole-codebase analysis, Qwen remains ahead.
License
GLM (Apache-like) ≈ Qwen (Apache 2.0 for most variants) ≈ Gemma 4 (Apache 2.0 since April 2026). Three straightforward licenses for enterprise use, with Gemma having dropped its custom clause.
Ecosystem
Qwen > Gemma 4 > GLM 5.1. Qwen now has the largest body of community fine-tunes, Gemma 4 is close behind, while GLM 5.1 remains more niche—fewer variants, fewer French tutorials.
i
Our honest recommendation
If you have 12 GB of VRAM and want to try something off the beaten path: GLM 5.1 9B. If you want the best open generalist at 24 GB: Qwen 3.8 27B or GLM 5.1 32B, depending on which runs better on your prompts. If community support and fine-tuning matter: Qwen 3.5 remains unbeatable in the number of derived options.

#Tips and pitfalls

Short system prompt
GLM 5.1 follows system prompts well but gets a bit saturated if you give it 800 words. Aim for 200–400 words, clear and structured.
Low temperature for FR
0.2–0.4 to minimize anglicisms and agreement oddities. Above 0.7, the quality of the French writing starts to fluctuate.
Stop tokens
The model uses <|endoftext|> and <|user|> as delimiters. If you write your own client, add these two strings to the stop tokens to prevent it from continuing to talk on its own.
Streaming via Ollama
The /api/chat endpoint and /v1/chat/completions support stream=true. There is no particular quirk with GLM 5.1, unlike some MoE models.
LoRA fine-tuning
GLM 5.1 9B is viable with LoRA on a RTX 4090 24 GB using Unsloth. The 32B requires multi-GPU or a Mac with generous unified memory.
Ollama Modelfile
To lock in a system prompt and parameters, create a Modelfile based on glm5.1:9b, then ollama create monassistant -f Modelfile. The derived model appears in ollama list like any other.
!
Confusion with ChatGLM and GLM-4
Don't confuse GLM 5.1 (a modern dense transformer architecture, released in 2026) with ChatGLM (the first generation, 2023) or GLM-4 (2024). The Ollama tags are distinct (chatglm, glm4, glm5.1); make sure you check what you're pulling.

#Go further

GLM 5.1 is running, and you have your first results. A few ways to push it further:

Choose your quantization (Q4, Q5, Q8, FP16)
For precisely deciding the quality/VRAM tradeoff for your use cases.
Qwen3.6 35B-A3B locally: testing and VRAM requirements
The direct competitor on the Qwen side, using an MoE architecture. Useful comparison if you’re undecided.
Open WebUI with Ollama: complete guide
A ChatGPT-like interface on top of GLM 5.1: conversations, RAG, and multiple users.
Build a local AI agent in Python with LangChain and Ollama
To use GLM 5.1's native tool calling and have it call your own functions.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.