Home › Reviews & tests › GLM Coding Plan
GLM Coding Plan: our review after testing
API access to GLM models, designed for coding assistants—by the lab that releases its weights as open weights.
The best value right now for connecting a coding assistant to a large model—precisely because the weights are open and you retain a local fallback. The caveat is real: some subscribers consider the plan quotas opaque.

What exactly is it?
GLM is the model family from the Chinese laboratory Zhipu AI (international brand: Z.ai), listed in Hong Kong since January 2026. Its distinctive feature, and the reason this site has been covering it for months, is that the weights are released as open weights. GLM-5, GLM-5.1, and GLM-5.2 are listed in our catalog, along with their VRAM requirements.
The Coding Plan is the code-focused API subscription: it exposes GLM models through a protocol compatible with existing assistants (Claude Code, Cline, Roo Code), at a significantly lower price than equivalent American APIs. The idea is simple: your coding tool can't tell the difference, but your bill can.
Our test: where local GLM stops on a consumer GPU
Before evaluating the API, we pushed the local option to its limit on our test machine (RTX 5070 Ti 12 GB + Intel Core Ultra 9 275HX, CachyOS, Ollama). That's the whole point of an open-weights family: the question isn't “API or nothing,” but “at what point does the API become necessary?” The answer is measured, with numbers to back it up:
| Measurement (26/08/2026) | Result |
|---|---|
| Model tested locally | GLM-4.7-flash (q4 quantized) |
| Weights in memory | 20 GB — for 12 GB of VRAM: 47% CPU / 53% GPU offloading |
| Generation (Python coding task) | 35 tokens/s — surprisingly smooth |
| Prompt processing (prefill) | 5.8 tokens/s — the real bottleneck |
| Loading the model | 25 s |
The measurements tell a more nuanced story than “it doesn't fit”: even when half-offloaded to the CPU, the model generates at 35 tokens/s — comfortable for chat. But the prompt processing tops out at 5.8 tokens/s: sending a 2,000-token file for analysis already takes more than five silent minutes before the first word. Yet a coding assistant spends its time doing exactly that—rereading your files on every iteration.
Unambiguous conclusion: on a 12 GB consumer GPU, this model family is unusable as a local coding agent, not generation but prefill. That's precisely the gap the Coding Plan fills: the same model family, served with instant prefill, while your GPU stays free. And if you have 24 GB of VRAM or more, run the numbers again— the configurator gives it to you for your machine.
Pricing
The Coding Plan comes in several tiers (from the entry-level plan for individual use to the “Max” tier for intensive multi-agent use). Prices observed in August 2026 range from a few dollars per month for the entry level to a few dozen dollars for the higher tiers—roughly an order of magnitude below equivalent Anthropic or OpenAI subscriptions. Z.ai frequently runs promotions on the first month: check the current price before subscribing, as it changes quickly.
Strengths and caveats
- Open weights = no lock-in: your prompts, your workflow, and even the model remain recoverable locally
- Compatible with existing coding assistants (Claude Code, Cline, Roo Code) without changing tools
- Starting price far below equivalent American APIs
- Solid company: Zhipu AI, listed in Hong Kong, one of the leading open-model labs
- Quotas considered opaque by some subscribers (“3× Claude” disputed during peak hours) — Trustpilot 2.1/5 from 32 reviews in August 2026; monitor your usage during the first week
- Slow customer support according to the same feedback
- Your requests pass through Z.ai's servers: exclude it for code under an NDA — this is true of every API, which is why local remains our reference
Frequently asked questions
Can GLM be used for free?
Yes. GLM models are released with open weights: quantized variants can be downloaded and run for free with Ollama or LM Studio, within the limits of your VRAM. The subscription only covers API access to the full variants, served on Z.ai's infrastructure.
Does the GLM Coding Plan work with Claude Code?
Yes, that's its main selling point: the API exposes a compatible protocol, so you just point the tool to Z.ai's endpoint. Cline and Roo Code are also supported. Configuration takes a few minutes.
Are the stated quotas reliable?
This is the main documented caveat. Some subscribers report that their quota is consumed faster than advertised during peak hours, and the service’s Trustpilot rating reflects this (2.1/5 from a small sample of 32 reviews). Our advice: take the entry tier, measure a week of real-world usage, then decide.
Where does my data go with this plan?
Your requests are processed on Z.ai's servers. For proprietary code under an NDA or regulated data, the rule remains the same: stay with a local model—this family makes that possible because the weights are open.
What machine do you need to run GLM locally?
The smaller quantized variants run on an 8 to 12 GB GPU. Recent MoE variants (GLM-4.7-flash: ~20 GB loaded in q4, measured on our machine) require more—you can measure this in our configurator, which tells you exactly what your machine can handle.
The other reviews in the selection
AI transcription, subtitling, and translation, with optional human review—the highest-rated European company in our selection.
Read the review →Turn a conversation into a reusable agent connected to your tools—without writing code.
Read the review →Online bootcamps—data science & AI, cybersecurity, UX, and web development—with German quality accreditation.
Read the review →Transparency. The “See GLM Coding Plan” link is an affiliate link (Impact.com platform). If you subscribe, QuelLLM.fr receives a commission paid by the brand, at no additional cost to you. This commission does not affect the rating or content: the free local alternative is mentioned before the purchase link, and caveats from customer reviews are published as-is. Complete methodology · Legal notices.