Advanced 14 minZhipu

GLM-5.2 locally: Ollama + LM Studio, the giant MIT

GLM-5.2 is the first frontier-scale open-weight model (753 billion parameters in a Mixture-of-Experts architecture) released under the MIT license and, above all, the only model in this category with confirmed GGUF quants tested on Ollama and LM Studio. Running GLM-5.2 locally is not a fantasy: it is possible on a Mac with generous unified memory or on a multi-GPU box, provided you accept aggressive quantization and modest speeds. This guide covers configurations that actually work, without overselling them.

By Mohamed Meguedmi·Update 2026-08-15·Tested on Windows, macOS, and Linux

#Why GLM-5.2 locally

Most “frontier” models (the largest and most capable) remain locked behind an API: GPT, Gemini, or the 100B+ versions of Qwen and DeepSeek. GLM-5.2 breaks that pattern. Zhipu AI published the full weights under the MIT license—the most permissive license there is, with no attribution or share-alike clause—and the community produced functional GGUF quants as soon as it was released. The result: you can host a frontier-class model at home, with no account, quota, or data leak.

The point isn’t speed—let’s be clear, a 753B model running locally will never match a cloud endpoint for responsiveness. The point is total sovereignty over a model that, in raw quality, competes with the best proprietary services: long-form reasoning, coding across large codebases, and a 1 million-token context. For a coding agent that runs for hours on proprietary code under an NDA, speed comes second to confidentiality.

i
In two words
GLM-5.2 = 753B MoE, MIT license, 1M context, with real GGUF quants on Ollama and LM Studio. It's the only frontier-size model we can honestly recommend self-hosting today—provided you have enough RAM or VRAM to run it.

#What has changed since GLM-5.1

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

If you're looking to install a reasonable GLM on a 12 to 24 GB GPU, this guide isn't the right one: GLM-5.1 (dense 9B and 32B) is designed for that, and our dedicated guide covers it in detail. GLM-5.2 targets a different audience. The confusion is common because the names are sequential, but the two models are in completely different hardware categories.

Architecture
GLM-5.1 is dense (9B, 32B). GLM-5.2 is an MoE with 753B total, with a fraction of experts activated per token. It cannot be deployed on a single consumer GPU.
License
GLM-5.1 is under the GLM License (an Apache variant with attribution). GLM-5.2 switches to pure MIT—commercial use with no restrictions and no viral clause.
Context
128k on GLM-5.1 (1M via degraded YaRN). GLM-5.2 natively handles 1M tokens, a real advantage for analyzing large repositories.
Hardware target
GLM-5.1: an 8–24 GB GPU. GLM-5.2: a Mac with 256 GB of unified memory, or a multi-GPU box, or a server with lots of RAM and offload.
Use cases
GLM-5.1 for a responsive local assistant. GLM-5.2 for a reference frontier model when quality matters more than speed.
→
Which one to choose
If your hardware tops out at 24 GB of VRAM, stick with GLM-5.1 32B: it will be faster and more comfortable. GLM-5.2 only makes sense if you have 128 GB or more of memory (unified or RAM+VRAM) and accept 5 to 15 tok/s in exchange for frontier quality.

#753B MoE: understanding the model

GLM-5.2 is a Mixture-of-Experts model: of its 753 billion parameters, only a fraction is activated for each generated token. That makes local inference feasible—the computation per token remains reasonable—but the catch is elsewhere: all the weights must fit in memory, including experts that are not being used at a given moment. Memory, not compute power, is the limiting factor.

Total parameters
753B, split between a shared backbone and a pool of dynamically routed experts.
Active parameters
A fraction per token (top-k routing). That is what enables decent throughput despite the total size.
Context
Native 1M tokens. Warning: a 1M KV cache consumes an enormous amount of memory; reserve it for cases that justify it.
License
MIT. You can fine-tune, redistribute, and integrate it into a commercial product without obligation.
Format
Original size in BF16 (~1.5 TB). Unusable as-is locally: this is where GGUF quants come in.
!
Memory first
Don't reason about it like a dense model. A 753B MoE activates only some of its weights per token, but you still have to load ALL of them into memory. A 2-bit quant brings the model down to about 200 GB; that is the figure that determines whether your machine can accommodate it, not the number of active parameters.

#GGUF quants (unsloth)

The community produced the quantizations that make GLM-5.2 usable. Unsloth's dynamic quants are the reference: they apply variable precision based on layer importance, preserving quality better than uniform quantization at the same weight. This is decisive at very low precision, where every bit counts.

Q2_K_XL (unsloth)
~200 GB. The realistic entry point. Degraded but usable quality; this is the version that fits on a 256 GB Mac or a multi-4090 box.
Q4_K_M
~380–400 GB. The quality sweet spot, but reserved for servers with very large memory or substantial multi-GPU setups.
Q5_K_M
~480 GB. Little noticeable gain with Q4 for this type of model; rarely justified locally.
Q8_0 / BF16
800 GB to 1.5 TB. The domain of professional GPU servers, beyond the reach of a consumer machine.
Download an unsloth quant (Hugging Face)
# huggingface-cli doit être installé : pip install -U huggingface_hub
# Q2_K_XL est réparti en plusieurs shards GGUF
huggingface-cli download unsloth/GLM-5.2-GGUF \
  --include "*Q2_K_XL*" \
  --local-dir ./glm-5.2-gguf
→
Why dynamic quants
At 2-bit, uniform quantization destroys model coherence. Unsloth dynamic quants keep sensitive layers (attention, embeddings) at higher precision and compress the rest aggressively. That’s what keeps a Q2_K_XL coherent where a naive Q2 goes off the rails.

#Realistic hardware requirements

There is no “lightweight” configuration for GLM-5.2. Here are the two profiles that actually work, without fudging the numbers.

Mac Apple Silicon 256 GB
A Mac Studio with an M-series chip and 256 GB of unified memory can run Q2_K_XL with room for context. Unified memory is a decisive advantage here: there’s no CPU/GPU split.
Mac 192 GB
Usable but tight: Q2_K_XL fits, but reduce the context window and close everything else. 128 GB is below the practical threshold.
Multi-GPU box
Several RTX 4090 cards (24 GB each) plus plenty of system RAM for CPU offloading. You can't fit 200 GB in VRAM alone; you split it between GPU and RAM.
Storage
At least 200 GB free for Q2, plus a fast NVMe SSD (initial loading reads hundreds of GB).
System RAM (box)
At least 128 GB of RAM if you offload experts to the CPU; 256 GB for a comfortable setup.

#1. Installation with LM Studio

LM Studio is often the simplest option for a model this size, especially on Mac: its MLX engine and memory management are well established, and the interface shows in real time how much memory the model needs before loading it. That is invaluable when you are pushing the limit.

  1. 01
    1. Install LM Studio
    Download LM Studio from the official website (lmstudio.ai) and install it. On a Mac, choose the native Apple Silicon version.
  2. 02
    2. Find the model
    In the search tab, type “GLM-5.2” and find the unsloth GGUF repository. LM Studio shows whether each quantization is compatible with your available RAM (green/orange/red badge).
  3. 03
    3. Choose the Q2_K_XL quantization
    Select the Q2_K_XL variant. LM Studio downloads all shards automatically—allow plenty of time depending on your connection (200 GB).
  4. 04
    4. Adjust the context
    Before loading, reduce the context length to a reasonable value (8k–16k for testing). Do not provide 1M tokens of input: the KV cache would cause memory usage to explode.
  5. 05
    5. Load and test
    Click “Load.” Monitor the memory gauge. Once loaded, send a first prompt and measure the displayed throughput in tok/s.
i
MLX vs GGUF on Mac
LM Studio also offers MLX versions (format Apple) for certain models. For GLM-5.2, the unsloth GGUF remains the confirmed and best-documented option. If a quantized MLX version appears, it may provide a slight speed boost on Apple Silicon, but first verify that it actually exists before looking for it.

#2. Installation with Ollama

Ollama can also serve GLM-5.2 from a GGUF, using a Modelfile that points to the downloaded files. This is the preferred route if you want to expose the model through an OpenAI-compatible API to other tools (agents, IDEs, Open WebUI).

Modelfile for a local GGUF
# Fichier : Modelfile
FROM ./glm-5.2-gguf/GLM-5.2-Q2_K_XL-00001-of-00005.gguf

PARAMETER num_ctx 16384
PARAMETER temperature 0.6
Create and run the model
# Créer l'entrée Ollama à partir du Modelfile
ollama create glm-5.2 -f Modelfile

# Lancer
ollama run glm-5.2

Once created, GLM-5.2 is served like any Ollama model at the default endpoint http://localhost:11434. Any OpenAI-compatible client can then query it by pointing to this address.

Check the GPU/CPU split
# Dans un autre terminal, après le premier prompt
ollama ps
!
Expected overflow
With a model this size, ollama ps will almost always show a GPU/CPU mix unless you have an overprovisioned machine. That's normal here, unlike with smaller models: the goal isn't 100% GPU usage, but fitting the model without disk swapping, which would seriously destroy performance.

#3. 256 GB Mac configuration

This is the most elegant configuration for GLM-5.2. Apple Silicon's unified memory means the GPU and CPU share the same pool: no costly transfers, no manual splitting. A Mac Studio with 256 GB loads Q2_K_XL and leaves room for a comfortable context.

Model loaded
Q2_K_XL (~200 GB) fits, with ~40-50 GB left for the KV cache and the system.
Expected throughput
Around 5 to 12 tok/s during generation, depending on context length. Comfortable for asynchronous work, frustrating for fast interactive chat.
Usable context
32k to 64k without a problem. Moving up to 128k+ is possible, but quickly eats into memory through the KV cache.
Recommended tool
LM Studio for simplicity, or Ollama if you connect agents to it.
→
Free up the GPU memory limit on Mac
By default, macOS reserves some unified memory for the system. To leave more RAM for the GPU on a high-end configuration, you can adjust iogpu.wired_limit_mb via sysctl. Do this carefully and test it: reserving too little for the system makes the machine unstable.

#4. RTX 4090 Box in 2-bit

On a PC, you have to deal with the VRAM/RAM split. A single RTX 4090 (24 GB) obviously cannot hold 200 GB: the strategy is to load as many experts as possible into VRAM and offload the rest to system RAM through llama.cpp. Throughput then depends directly on the GPU/CPU ratio and your RAM speed.

Launching llama.cpp with offload
./llama-server \
  -m glm-5.2-gguf/GLM-5.2-Q2_K_XL-00001-of-00005.gguf \
  -c 16384 \
  -ngl 99 \
  --n-cpu-moe 40 \
  -fa \
  --host 0.0.0.0 --port 8080
-ngl 99
Attempts to place as many layers as possible on the GPU. llama.cpp fills the available VRAM and moves the rest to the CPU.
--n-cpu-moe 40
Keep the indicated MoE layers on the CPU. This is the key lever for an MoE model: put the dense backbone in VRAM and offload the experts to RAM. Adjust the number based on your VRAM.
-fa
FlashAttention reduces KV-cache memory usage. Keep it enabled.
Multi-GPU
With 2 to 4 RTX 4090, llama.cpp automatically distributes the workload (--split-mode). More VRAM = less CPU offloading = better throughput.
!
Throughput drops quickly with offloading
Every layer offloaded to the CPU is costly. On a single 4090 with most experts in RAM, expect 2 to 5 tok/s — it’s slow. Multi-GPU improves the situation significantly. On a PC, GLM-5.2 is an exercise in patience; reserve it for tasks where the quality justifies the wait.

#5. Long-running coding agents

This is where GLM-5.2 running locally makes perfect sense despite its slowness. A coding agent that refactors a large codebase, reads dozens of files, and reasons for hours does not care about a latency of a few seconds per token: it works in the background. What matters is reasoning quality and the 1M context that lets it keep the entire repository in mind, all without a line of proprietary code leaving your machine.

Massive context
A million tokens lets you inject an entire codebase into the prompt instead of fragmenting it, improving the consistency of multi-file changes.
Tool calling
GLM-5.2 handles tool calls in the OpenAI format, which is essential for an agent that reads, writes, and executes.
OpenAI endpoint
Through Ollama (localhost:11434) or llama.cpp (localhost:8080), connect Aider, Cline, or any agent that speaks the OpenAI API.
Asynchronous work
Start the task and do something else. At 5–10 tok/s, a large refactor takes time but runs without supervision.
i
Privacy as the main argument
For code covered by an NDA or industrial secret, local GLM-5.2 is one of the few ways to achieve near-frontier quality without ever transmitting the code to a third party. Slowness becomes an acceptable trade-off compared with the legal and leakage risk involved in using a cloud agent.

#Honest expectations and troubleshooting

No serious guide will claim that running a 753B locally is smooth. Here are the real problems and their fixes.

Disk swap = death
If the model spills into disk swap, throughput drops to seconds per token. Reduce the context, choose a smaller quantization, or close other RAM-hungry applications.
Very slow loading
Loading 200 GB from an SSD takes several minutes. That's normal. A fast NVMe SSD makes a real difference; an external USB drive should be avoided.
OOM during loading
On Mac, adjust the GPU limit (iogpu.wired_limit_mb). On PC, increase --n-cpu-moe to offload more experts to RAM.
Lower quality
At 2-bit, the model may hallucinate more. Lower the temperature (0.6 or less) and favor unsloth dynamic quants over a uniform Q2.
Disappointing throughput
This is expected. A frontier-size model running locally is not fast. If you want speed, GLM-5.1 32B or a Qwen3 will be much more responsive.

#Go further

GLM-5.2 is an extreme case involving quantization, hardware, and agent deployment. These guides cover the fundamentals you need to master around it.

The sensible little sibling
“GLM 5.1 locally: the open-weight alternative to know about”—if your hardware tops out at 24 GB, GLM is the model you need, and it's much more responsive.
Understanding quants
“Choosing your quantization (Q4, Q5, Q8, FP16)” — essential for understanding why dynamic 2-bit makes GLM-5.2 accessible without destroying it.
Coding with a local agent
“Aider + Ollama: coding in the terminal with a 100% local agent” — the starting point for connecting GLM-5.2 to a real development workflow.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.