Beginner 12 minOverview

Local Chinese LLM: Qwen, DeepSeek, GLM, Kimi

In 2026, if you download an open-weight LLM to run at home, there’s a good chance it will come from China. Qwen, DeepSeek, GLM, Kimi: these families dominate the top of the open-model rankings, often under more permissive licenses than those from Western labs. This overview sorts it out: what each family is really worth, what its licenses actually say, and why running a Chinese LLM locally answers the privacy question upfront.

By Mohamed Meguedmi·Update 2026-09-01·Tested on Windows, macOS, and Linux

#Why Chinese LLMs dominate open-weight models

The open-model landscape has shifted. Where Meta (Llama) was leading the race two years ago, Chinese labs now publish the highest-performing and most frequent releases. Alibaba (Qwen), DeepSeek, Zhipu AI (GLM), and Moonshot (Kimi) release models at a steady pace, from the small 3B to MoE models with several hundred billion parameters.

The reason is as much strategic as technical: publishing the weights as open-weight models helps them gain global visibility, attract the community, and become de facto standards, including against closed American models. For you, as a local user, the result is simple: an immense selection of free, regularly updated models that often go toe-to-toe with the cloud on coding, reasoning, and multilingual capabilities.

i
Open-weight ≠ open-source
“Open-weight” means that the model weights can be downloaded and run freely. It does not guarantee that the training data or complete code has been published. That is the case for nearly all the models in this guide—you get the model, not its recipe.

#Local deployment solves the privacy issue

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the knee-jerk objection: “a Chinese LLM means my data goes to China.” That's true for cloud APIs (chat.qwen.ai, the DeepSeek app, the Zhipu platform)—your prompts do pass through remote servers subject to Chinese law. But this guide is about local use, and locally that concern simply no longer applies.

When you run the model with Ollama or LM Studio, you download a fixed weight file (a GGUF, for example), and your GPU calculates the responses. The model has no network capability: it isn't software that “phones home”; it's a large matrix of numbers. Nothing leaves your machine, regardless of where the model comes from.

In the cloud (API)
Your prompts go to the vendor. The model’s origin and jurisdiction really matter.
Locally
The weights run on your hardware, offline if you want. No data leaves the machine.
The actual residual risk
A bias or censorship in the model’s responses (on certain sensitive topics), not a data leak. This is an output-quality issue, not a privacy issue.
→
Verify that nothing gets out
To convince yourself, launch Ollama, turn off Wi-Fi, and chat with the model. Everything keeps working. A local LLM only needs the network for the initial weight download.

#The actual licenses: Apache 2.0 and MIT

Another misconception: “Chinese licenses are opaque or full of traps.” The reality in 2026 is the opposite—they’re often the cleanest licenses on the market. Most models in this overview are released under Apache 2.0 or MIT, two permissive international licenses, including for commercial use.

Qwen
Most recent variants, including Qwen3.8-27B, are under Apache 2.0—free commercial use, redistribution allowed.
DeepSeek
The V3/V4 and R1 models are released under the MIT license, one of the most permissive licenses available.
GLM (Zhipu)
The latest generations, including GLM-5.2, have moved to the MIT license.
Kimi (Moonshot)
Open weights under a permissive MIT/Apache-type license, depending on the version.
!
Still, read the model card
Some older families or specific variants may carry a custom license with usage restrictions. Before a commercial deployment, always check the LICENSE file on the Hugging Face page for the specific model you're downloading—not just the family's reputation.

#Qwen (Alibaba): the Swiss Army knife

Qwen is the most complete and versatile family in the ecosystem. Alibaba offers the model in every size, from the 3B that fits on an entry-level GPU to large MoEs, and for every specialty: general-purpose, code (Qwen3-Coder), vision (Qwen3-VL), and a “thinking” mode for reasoning.

Strengths
Excellent multilingual performance (including very good French), an unmatched range of sizes, a mature tooling ecosystem, and day-0 availability on Ollama.
Typical sizes
From 3B/7B models for testing, to 27B (Qwen3.8) as the quality/VRAM sweet spot, and up to 35B-A3B MoE models for high-end GPUs.
Who it's for
The first default choice if you are just getting started and want a model that handles everything correctly.

#DeepSeek: the reasoning champion

DeepSeek became known with R1, a chain-of-thought reasoning model that shook up the industry in early 2025. The family remains the benchmark for logical, mathematical, and complex coding tasks. Distilled versions of R1 (7B, 14B, 32B) make this reasoning accessible on consumer hardware.

Strengths
Very high-level reasoning and mathematics, MIT license, distilled versions playable locally.
The catch
The flagship models (V3/V4, 671B and larger in MoE) are enormous—beyond the reach of a single GPU. For consumer-grade local use, target the R1 distillations.
Who it's for
People doing difficult coding, math, or logic—or who want to see the model “think” before responding.

#GLM (Zhipu): the versatile outsider

GLM, developed by Zhipu AI, is the serious challenger to Qwen. It performs well in French and is strong at coding and reasoning, offering reasonable sizes (around 9B and 32B) aimed directly at the sweet spot for 12–24 GB GPUs, as well as giant MoE models such as GLM-5.2 (753B) for extreme configurations.

Strengths
Excellent quality-to-size balance across the 9B–32B variants, good French, MIT license on the latest generations.
Typical sizes
9B for an 8–12 GB GPU, 32B for 24 GB. Frontier MoE versions remain limited to Macs with large amounts of unified memory.
Who it's for
A credible alternative to Qwen when you want to compare, or a model that performs well in French on modest hardware.

#Kimi and MiniCPM: the two extremes

These two families illustrate the ends of the spectrum. Kimi (Moonshot) targets the very large end: MoE models with 1 trillion parameters and beyond, known for their long context and agentic quality. MiniCPM (the OpenBMB / ModelBest team) targets the opposite: compact, efficient models designed to run on mobile devices or tiny GPUs.

Kimi K2 / K3
Giant MoEs (1T to 2.8T parameters). Open weights, but “open” does not mean “runnable on your machine”: they require dozens of accelerators. Reserved for servers, not workstations.
MiniCPM
Ultra-compact models, some multimodal, designed for embedded systems and small configurations. Ideal when every gigabyte of VRAM matters.
Key takeaway
Open weights cover the entire hardware spectrum. Don't confuse “the model is free” with “my machine can run it.”

#Which one to choose based on your VRAM

The real limiting factor locally isn't the model brand but your GPU's VRAM. With Q4_K_M quantization (the best quality-to-size compromise), here are concrete guidelines for grouping these families by hardware tier.

8 GB (RTX 3060 8G, 4060)
Qwen 7B, GLM 9B, DeepSeek-R1 distilled 7B in Q4 (~5 GB). Comfortable for chat and light coding.
12 GB (RTX 3060 12G, 4070)
The sweet spot. Qwen 14B, GLM 9B in Q8, DeepSeek-R1 14B (~9 GB). Room for a larger context.
16 GB (RTX 4080, 5070 Ti)
Qwen ~24-27B, DeepSeek-R1 32B in a tight quantization. The serious tier for reasoning.
24 GB (RTX 3090/4090)
Qwen 27–32B and GLM 32B entirely in VRAM (~19 GB), MoE 35B-A3B. The best of mainstream local models.
Mac with unified memory (48–128+ GB)
The only realistic option for large MoE models (GLM-5.2, DeepSeek Flash) via offloading to unified memory—at moderate throughput.
i
VRAM benchmark for Q4
3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. Always add 1 to 2 GB for context and the system. A model that exceeds VRAM spills into RAM (CPU offload), and its throughput collapses.

#Installing a Chinese model in practice

The typical stack is identical regardless of the model’s origin: Ollama as the inference daemon (listening on http://localhost:11434), with Open WebUI or LM Studio as the interface. Here’s how to obtain a model from each family.

  1. 01
    Install Ollama
    Download the daemon from ollama.com and launch it. It exposes a local API on port 11434 — nothing else needs to be configured to get started.
  2. 02
    Choose a model based on your VRAM
    Refer to the table above. If in doubt, go with Qwen 14B or GLM 9B in Q4_K_M: they fit on most recent GPUs and cover 90% of use cases.
  3. 03
    Download and run
    A single command downloads the weights and opens the chat. The first launch retrieves several gigabytes; subsequent launches are instant.
  4. 04
    Check offline mode
    Once the weights are cached, disconnect from the network and keep chatting: that proves the model is running 100% locally.
Terminal
# Qwen généraliste (Alibaba, Apache 2.0)
ollama run qwen3

# DeepSeek-R1 distillé, raisonnement (MIT)
ollama run deepseek-r1:14b

# GLM, l'alternative polyvalente (Zhipu)
ollama run glm4

# Vérifier les modèles installés et que le daemon répond
curl http://localhost:11434/api/tags
→
Tag names vary
The exact tags (qwen3, deepseek-r1:14b, etc.) and the available versions change quickly in the Ollama library. Check ollama.com/library for the exact name and size in GB of each variant before downloading.

#Go further

This overview gives you the map; these guides go into the details of each family and local settings:

The Qwen sweet spot
“Qwen 3.8 27B locally: Ollama installation, VRAM, and settings” details the exact command, VRAM by quantization, and thinking mode.
The DeepSeek reasoning
“Installing DeepSeek R1 with Ollama” covers the distilled 7B/14B/32B versions and local chain-of-thought.
Choose based on your card
The “Which LLM for 12 GB / 16 GB / 24 GB of VRAM?” guides turn these tiers into concrete GPU recommendations.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.