GLM-5.2 locally: Ollama + LM Studio, the giant MIT
GLM-5.2 is the first frontier-scale open-weight model (753 billion parameters in a Mixture-of-Experts architecture) released under the MIT license and, above all, the only model in this category with confirmed GGUF quants tested on Ollama and LM Studio. Running GLM-5.2 locally is not a fantasy: it is possible on a Mac with generous unified memory or on a multi-GPU box, provided you accept aggressive quantization and modest speeds. This guide covers configurations that actually work, without overselling them.
#Why GLM-5.2 locally
Most “frontier” models (the largest and most capable) remain locked behind an API: GPT, Gemini, or the 100B+ versions of Qwen and DeepSeek. GLM-5.2 breaks that pattern. Zhipu AI published the full weights under the MIT license—the most permissive license there is, with no attribution or share-alike clause—and the community produced functional GGUF quants as soon as it was released. The result: you can host a frontier-class model at home, with no account, quota, or data leak.
The point isn’t speed—let’s be clear, a 753B model running locally will never match a cloud endpoint for responsiveness. The point is total sovereignty over a model that, in raw quality, competes with the best proprietary services: long-form reasoning, coding across large codebases, and a 1 million-token context. For a coding agent that runs for hours on proprietary code under an NDA, speed comes second to confidentiality.
#What has changed since GLM-5.1
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
If you're looking to install a reasonable GLM on a 12 to 24 GB GPU, this guide isn't the right one: GLM-5.1 (dense 9B and 32B) is designed for that, and our dedicated guide covers it in detail. GLM-5.2 targets a different audience. The confusion is common because the names are sequential, but the two models are in completely different hardware categories.
- Architecture
- GLM-5.1 is dense (9B, 32B). GLM-5.2 is an MoE with 753B total, with a fraction of experts activated per token. It cannot be deployed on a single consumer GPU.
- License
- GLM-5.1 is under the GLM License (an Apache variant with attribution). GLM-5.2 switches to pure MIT—commercial use with no restrictions and no viral clause.
- Context
- 128k on GLM-5.1 (1M via degraded YaRN). GLM-5.2 natively handles 1M tokens, a real advantage for analyzing large repositories.
- Hardware target
- GLM-5.1: an 8–24 GB GPU. GLM-5.2: a Mac with 256 GB of unified memory, or a multi-GPU box, or a server with lots of RAM and offload.
- Use cases
- GLM-5.1 for a responsive local assistant. GLM-5.2 for a reference frontier model when quality matters more than speed.
#753B MoE: understanding the model
GLM-5.2 is a Mixture-of-Experts model: of its 753 billion parameters, only a fraction is activated for each generated token. That makes local inference feasible—the computation per token remains reasonable—but the catch is elsewhere: all the weights must fit in memory, including experts that are not being used at a given moment. Memory, not compute power, is the limiting factor.
- Total parameters
- 753B, split between a shared backbone and a pool of dynamically routed experts.
- Active parameters
- A fraction per token (top-k routing). That is what enables decent throughput despite the total size.
- Context
- Native 1M tokens. Warning: a 1M KV cache consumes an enormous amount of memory; reserve it for cases that justify it.
- License
- MIT. You can fine-tune, redistribute, and integrate it into a commercial product without obligation.
- Format
- Original size in BF16 (~1.5 TB). Unusable as-is locally: this is where GGUF quants come in.
#GGUF quants (unsloth)
The community produced the quantizations that make GLM-5.2 usable. Unsloth's dynamic quants are the reference: they apply variable precision based on layer importance, preserving quality better than uniform quantization at the same weight. This is decisive at very low precision, where every bit counts.
- Q2_K_XL (unsloth)
- ~200 GB. The realistic entry point. Degraded but usable quality; this is the version that fits on a 256 GB Mac or a multi-4090 box.
- Q4_K_M
- ~380–400 GB. The quality sweet spot, but reserved for servers with very large memory or substantial multi-GPU setups.
- Q5_K_M
- ~480 GB. Little noticeable gain with Q4 for this type of model; rarely justified locally.
- Q8_0 / BF16
- 800 GB to 1.5 TB. The domain of professional GPU servers, beyond the reach of a consumer machine.
#Realistic hardware requirements
There is no “lightweight” configuration for GLM-5.2. Here are the two profiles that actually work, without fudging the numbers.
- Mac Apple Silicon 256 GB
- A Mac Studio with an M-series chip and 256 GB of unified memory can run Q2_K_XL with room for context. Unified memory is a decisive advantage here: there’s no CPU/GPU split.
- Mac 192 GB
- Usable but tight: Q2_K_XL fits, but reduce the context window and close everything else. 128 GB is below the practical threshold.
- Multi-GPU box
- Several RTX 4090 cards (24 GB each) plus plenty of system RAM for CPU offloading. You can't fit 200 GB in VRAM alone; you split it between GPU and RAM.
- Storage
- At least 200 GB free for Q2, plus a fast NVMe SSD (initial loading reads hundreds of GB).
- System RAM (box)
- At least 128 GB of RAM if you offload experts to the CPU; 256 GB for a comfortable setup.
#1. Installation with LM Studio
LM Studio is often the simplest option for a model this size, especially on Mac: its MLX engine and memory management are well established, and the interface shows in real time how much memory the model needs before loading it. That is invaluable when you are pushing the limit.
- 011. Install LM StudioDownload LM Studio from the official website (lmstudio.ai) and install it. On a Mac, choose the native Apple Silicon version.
- 022. Find the modelIn the search tab, type “GLM-5.2” and find the unsloth GGUF repository. LM Studio shows whether each quantization is compatible with your available RAM (green/orange/red badge).
- 033. Choose the Q2_K_XL quantizationSelect the Q2_K_XL variant. LM Studio downloads all shards automatically—allow plenty of time depending on your connection (200 GB).
- 044. Adjust the contextBefore loading, reduce the context length to a reasonable value (8k–16k for testing). Do not provide 1M tokens of input: the KV cache would cause memory usage to explode.
- 055. Load and testClick “Load.” Monitor the memory gauge. Once loaded, send a first prompt and measure the displayed throughput in tok/s.
#2. Installation with Ollama
Ollama can also serve GLM-5.2 from a GGUF, using a Modelfile that points to the downloaded files. This is the preferred route if you want to expose the model through an OpenAI-compatible API to other tools (agents, IDEs, Open WebUI).
Once created, GLM-5.2 is served like any Ollama model at the default endpoint http://localhost:11434. Any OpenAI-compatible client can then query it by pointing to this address.
#3. 256 GB Mac configuration
This is the most elegant configuration for GLM-5.2. Apple Silicon's unified memory means the GPU and CPU share the same pool: no costly transfers, no manual splitting. A Mac Studio with 256 GB loads Q2_K_XL and leaves room for a comfortable context.
- Model loaded
- Q2_K_XL (~200 GB) fits, with ~40-50 GB left for the KV cache and the system.
- Expected throughput
- Around 5 to 12 tok/s during generation, depending on context length. Comfortable for asynchronous work, frustrating for fast interactive chat.
- Usable context
- 32k to 64k without a problem. Moving up to 128k+ is possible, but quickly eats into memory through the KV cache.
- Recommended tool
- LM Studio for simplicity, or Ollama if you connect agents to it.
#4. RTX 4090 Box in 2-bit
On a PC, you have to deal with the VRAM/RAM split. A single RTX 4090 (24 GB) obviously cannot hold 200 GB: the strategy is to load as many experts as possible into VRAM and offload the rest to system RAM through llama.cpp. Throughput then depends directly on the GPU/CPU ratio and your RAM speed.
- -ngl 99
- Attempts to place as many layers as possible on the GPU. llama.cpp fills the available VRAM and moves the rest to the CPU.
- --n-cpu-moe 40
- Keep the indicated MoE layers on the CPU. This is the key lever for an MoE model: put the dense backbone in VRAM and offload the experts to RAM. Adjust the number based on your VRAM.
- -fa
- FlashAttention reduces KV-cache memory usage. Keep it enabled.
- Multi-GPU
- With 2 to 4 RTX 4090, llama.cpp automatically distributes the workload (--split-mode). More VRAM = less CPU offloading = better throughput.
#5. Long-running coding agents
This is where GLM-5.2 running locally makes perfect sense despite its slowness. A coding agent that refactors a large codebase, reads dozens of files, and reasons for hours does not care about a latency of a few seconds per token: it works in the background. What matters is reasoning quality and the 1M context that lets it keep the entire repository in mind, all without a line of proprietary code leaving your machine.
- Massive context
- A million tokens lets you inject an entire codebase into the prompt instead of fragmenting it, improving the consistency of multi-file changes.
- Tool calling
- GLM-5.2 handles tool calls in the OpenAI format, which is essential for an agent that reads, writes, and executes.
- OpenAI endpoint
- Through Ollama (localhost:11434) or llama.cpp (localhost:8080), connect Aider, Cline, or any agent that speaks the OpenAI API.
- Asynchronous work
- Start the task and do something else. At 5–10 tok/s, a large refactor takes time but runs without supervision.
#Honest expectations and troubleshooting
No serious guide will claim that running a 753B locally is smooth. Here are the real problems and their fixes.
- Disk swap = death
- If the model spills into disk swap, throughput drops to seconds per token. Reduce the context, choose a smaller quantization, or close other RAM-hungry applications.
- Very slow loading
- Loading 200 GB from an SSD takes several minutes. That's normal. A fast NVMe SSD makes a real difference; an external USB drive should be avoided.
- OOM during loading
- On Mac, adjust the GPU limit (iogpu.wired_limit_mb). On PC, increase --n-cpu-moe to offload more experts to RAM.
- Lower quality
- At 2-bit, the model may hallucinate more. Lower the temperature (0.6 or less) and favor unsloth dynamic quants over a uniform Q2.
- Disappointing throughput
- This is expected. A frontier-size model running locally is not fast. If you want speed, GLM-5.1 32B or a Qwen3 will be much more responsive.
#Go further
GLM-5.2 is an extreme case involving quantization, hardware, and agent deployment. These guides cover the fundamentals you need to master around it.
- The sensible little sibling
- “GLM 5.1 locally: the open-weight alternative to know about”—if your hardware tops out at 24 GB, GLM is the model you need, and it's much more responsive.
- Understanding quants
- “Choosing your quantization (Q4, Q5, Q8, FP16)” — essential for understanding why dynamic 2-bit makes GLM-5.2 accessible without destroying it.
- Coding with a local agent
- “Aider + Ollama: coding in the terminal with a 100% local agent” — the starting point for connecting GLM-5.2 to a real development workflow.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.