BestLLMfor Your hardware. Your LLM. Your call.
The Local Copilot Kit APIOpen data Find my LLM

HomeTool Reviews › GLM Coding Plan

GLM Coding Plan: our verdict after benchmarking the local option first

API access to Z.ai’s GLM models, built for coding agents — from the lab that still publishes its weights.

Updated August 28, 2026 · by Mohamed Meguedmi · Contains affiliate links

4.0
/ 5
Our verdict

The best value we’ve found for wiring a coding agent to a large model, precisely because the weights stay open and you always keep a local exit ramp. The real reserve: a meaningful slice of subscribers call the weekly quota opaque, and Trustpilot reflects that.

Good fit forDevelopers running Claude Code, Cline, or Roo Code who want a large coding model without US-API pricing — and local-GLM users who keep hitting a VRAM wall.
Skip it ifCode under a strict NDA (requests transit Z.ai’s servers — stay local for that); anyone who needs a contractually guaranteed quota rather than a best-effort one.

What is it, exactly?

GLM is the model family from Zhipu AI (international brand: Z.ai), listed in Hong Kong since January 2026. What earns it a spot on this site: the weights ship open. GLM-5, GLM-5.1, and GLM-5.2 are all in our catalog, VRAM requirements included.

The Coding Plan is the API subscription aimed at coding workflows specifically: it exposes GLM through a protocol compatible with the agents developers already run — Claude Code, Cline, Roo Code — at a price well under the equivalent US APIs. Your coding tool doesn’t know the difference. Your bill does.

What we tested: where local GLM actually stops on a consumer GPU

Before judging the hosted plan, we pushed the free local option to its limit on our own rig: an RTX 5070 Ti (12 GB) paired with an Intel Core Ultra 9 275HX, running CachyOS and Ollama. That’s the whole point of an open-weights family — the question isn’t “API or nothing,” it’s “at what point does the API start paying for itself.” Measured August 26, 2026:

MeasurementResult
Model tested locallyGLM-4.7-flash (q4 quantization)
Weights in memory20 GB — against 12 GB of VRAM: 47% CPU / 53% GPU offload
Generation (Python coding task)35 tokens/sec — surprisingly usable
Prompt processing (prefill)5.8 tokens/sec — the real bottleneck
Model load time25 sec

The number that matters here isn’t generation speed — it’s prefill. Even split across CPU and GPU, GLM-4.7-flash generates at 35 tokens/sec, which is fine for a chat window. But feed it a 2,000-token file to review and you’re staring at five-plus minutes of silence before the first token, because prompt processing caps out at 5.8 tokens/sec. A coding agent spends most of its time doing exactly that — re-reading your files on every single turn.

Verdict: on a consumer 12 GB card, this specific model isn’t usable as a local coding agent, and the bottleneck is prefill, not generation. That’s precisely the gap the Coding Plan fills — the same model family, served with instant prompt processing, while your own GPU stays free. Sitting on 24 GB of VRAM or more? Redo the math for your card — our configurator does it in seconds.

Pricing

The Coding Plan runs three monthly tiers: Lite at $18/month, Pro at $80/month, and Max at $168/month, each with its own 5-hour and weekly prompt caps on a points-based quota. Annual billing cuts those to roughly $12.60 / $56 / $117.60 a month. Off-peak usage is billed at 50% of the standard rate — worth knowing if your workflow can tolerate running heavy jobs overnight.

Strengths and reserves

Strengths
  • Open weights = no lock-in: your workflow, and the model itself, stay recoverable locally
  • Drop-in compatible with the coding agents developers already use (Claude Code, Cline, Roo Code)
  • Entry price well below the equivalent US APIs, with a real off-peak discount
  • Backed by a real, publicly listed company — Zhipu AI is one of the major open-weight labs
Reserves
  • Quota transparency is the recurring complaint: Trustpilot sits at 2.1/5 on a small sample (32 reviews, August 2026), almost entirely about caps burning faster than advertised at peak hours — watch your usage the first week
  • Support is reported as slow on the same review threads
  • Requests transit Z.ai’s servers — rule out NDA code, same as with any hosted API, which is exactly why local stays our default recommendation
See GLM Coding Plan → Affiliate link — commission paid by the brand, at no extra cost to you

Frequently asked questions

Can I use GLM for free?

Yes. GLM models ship as open weights — quantized versions download and run for free with Ollama or LM Studio, within your VRAM budget. The subscription only covers API access to the full-size variants, served on Z.ai’s infrastructure.

Does the GLM Coding Plan work with Claude Code?

Yes, that’s its core pitch — the API exposes a compatible protocol, so you point your existing tool at Z.ai’s endpoint. Cline and Roo Code are supported too. Setup takes a few minutes.

Are the advertised quotas reliable?

That’s the main documented reserve. A share of subscribers report burning through their quota faster than advertised during peak hours, and Trustpilot reflects it (2.1/5 on a small 32-review sample). Our advice: start on Lite, measure a real week of usage, then decide.

Where does my code go on this plan?

Your requests are processed on Z.ai’s servers. For proprietary code under NDA or regulated data, the rule doesn’t change: stay local — which this specific model family makes genuinely possible, since the weights are open.

What hardware do I need to run GLM locally instead?

Smaller quantized variants fit on 8–12 GB GPUs. The newer MoE variants are heavier (GLM-4.7-flash: ~20 GB loaded at q4, measured on our own rig) — our configurator tells you exactly what your machine can handle.

The rest of the selection

HappyScribe4.5/5
Transcription & subtitles

AI transcription, subtitles, and translation, with optional human proofreading — the best-reviewed European brand in this lineup.

Read the review →
Creao AI3.5/5
AI agents

Turn a conversation into a reusable agent wired into your own tools — without writing code.

Read the review →
Code Labs Academy3.5/5
Education

Online bootcamps — data science & AI, cybersecurity, UX, web development — with an audited German quality accreditation.

Read the review →
Transkriptor4.3/5
Transcription & meeting notes

AI transcription and meeting notes, with automatic action-item extraction and calendar-triggered recording.

Read the review →
InVideo AI3.5/5
AI video generation

Prompt or script in, finished video out — voiceover, footage, captions, and music assembled automatically.

Read the review →

Transparency. The “See GLM Coding Plan” link is an affiliate link (Impact.com platform). If you subscribe, BestLLMfor earns a commission paid by the brand, at no extra cost to you. That commission influences neither the score nor the content: the free local alternative is cited before the buy link, and the reserves found in customer reviews are published as-is. Full methodology.