Distillation: how small models inherit gros
You’ve probably come across names like DeepSeek-R1-Distill-Qwen-14B and wondered why a “small” model suddenly reasons almost as well as a giant. The answer comes down to one word: distillation. It’s the technique that lets a compact model inherit the behavior of a much larger model without matching its size or slowness. This guide explains LLM distillation using the teacher-and-student analogy, details what is actually transferred and what gets lost along the way, and shows how to recognize successful distillation before running it on your own system.
#The problem distillation solves
The best open models are huge: several hundred billion parameters. They're brilliant, but unusable on a consumer machine: they require tens of gigabytes of VRAM and generate tokens slowly. By contrast, a 7B or 14B model fits on a gaming card and responds quickly, but it often lacks the larger model's reasoning finesse.
So we would like to recover some of a large model’s abilities in a format that runs locally. The naive approach would be to simply retrain a small model on the same raw data as the large one. That is not enough: at a smaller scale, learning alone from an ocean of text produces a decent model, but not an exceptional one. Distillation offers a shortcut: instead of relearning the world from scratch, the small model learns directly from the large one.
#The teacher-student principle
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Distillation involves two models. The “teacher” is the highly capable large model we want to imitate. The “student” is the small model we train. The principle is simple: ask the teacher thousands of questions, collect its answers, and train the student to produce the same answers. The student no longer learns from raw text found on the web, but from examples already “chewed up” by an expert model.
The subtlety that makes all the difference is exactly what the student copies. In the richest version, they do not imitate only the final word chosen by the teacher, but the entire probability distribution for the next word: the teacher does not merely say “the answer is cat,” but indicates “cat very likely, dog somewhat likely, toaster virtually impossible.” These nuances—sometimes called soft labels—contain information that the simple correct answer does not provide.
- Teacher (teacher)
- The large reference model, expensive but excellent. It serves as the source; it is not deployed on your system.
- Student (student)
- The small model trained to reproduce the professor’s behavior. This is the one you will run locally.
- Soft labels
- The teacher's nuanced probabilities for every possible word, not just the winning word. The richest raw material for distillation.
- Distillation dataset
- The full set of questions asked and the teacher's answers, which the student practices on.
#What is actually transmitted
People often imagine that distillation “transfers the knowledge” of the larger model. That's partly true, but imprecise. What transfers best is mainly behavior and ways of working, rather than raw facts. A small model has limited storage capacity: it cannot retain everything from a giant's internal encyclopedia. However, it can capture methods very well.
- Reasoning style
- How to break a problem into steps and formulate a chain of thought before reaching a conclusion. This is what transfers best, and it is the core of R1-Distill.
- Response format
- The structure, tone, and way of organizing an explanation or code. The student adopts the teacher's writing habits.
- Resolution schemes
- In math, code, or logic, the student inherits proven techniques: how to approach a type of problem and which checks to perform.
- Part of the knowledge
- Facts and associations, yes, but limited by the small model's capacity. This is the part it handles least well.
This distinction is essential for understanding recent distilled models. When DeepSeek distills R1 into a Qwen 14B, the goal is not to fit all of R1's general knowledge into 14 billion parameters—that is impossible. The goal is to transfer its way of reasoning: thinking step by step, correcting itself, and working through a proof. And that is something a small model can learn remarkably well.
#Why a distilled 14B beats a standard 14B
Here is the counterintuitive part. Two models are exactly the same size—14 billion parameters—so they use the same VRAM and run at the same speed. Yet the distilled model often crushes the classic one at reasoning. The size has not changed; only the training method has. How can the same number of parameters produce a much better model?
Because training doesn't fill parameters randomly: it organizes them. A standard 14B learns on its own from raw text; it develops solid but general-purpose capabilities. A distilled 14B learns from examples produced by an elite teacher, often with complete reasoning traces. Its 14 billion parameters are wired around good methods rather than a little of everything. At equal capacity, the quality of the training signal makes the difference.
- Same cost, better signal
- The distilled model is neither larger nor slower, but it was trained on much higher-quality material: an expert’s outputs.
- Some dense examples
- The professor's answers are full of correct and complete reasoning, rare in the web's raw text.
- Less noise
- The web contains plenty of errors and clutter. A teacher filters it: the student learns from clean material.
Be careful not to overinterpret this: “beating” applies mainly to the tasks targeted by distillation, typically reasoning, math, and code. On pure general knowledge or fine-grained factual knowledge, the gap narrows or even reverses. A distilled model is not magically superior at everything—it is superior where it was trained to be.
#The R1-Distill case
The example that popularized local distillation is the DeepSeek-R1-Distill family released in early 2025. DeepSeek used its large R1 reasoning model as the teacher and distilled its behavior into several students of different sizes, based on existing architectures such as Qwen and Llama. The result: 1.5B, 7B, 8B, 14B, and 32B models that reason with a visible chain of thought, something previously reserved for giants.
- R1-Distill-Qwen-1.5B
- Tiny, runs virtually anywhere, even without a GPU. Surprisingly capable for its size on simple math.
- R1-Distill-Qwen-7B / Llama-8B
- The mainstream compromise: ~5 GB in Q4, on an 8 GB card. Good reasoning for everyday use.
- R1-Distill-Qwen-14B
- ≈ 9 GB in Q4, the target for a RTX 3060 12 GB or 4070 12 GB. The sweet spot for reasoning and accessible hardware.
- R1-Distill-Qwen-32B
- ≈ 19 GB in Q4, for a RTX 4090 24 GB. The closest to the professor, if your VRAM can handle it.
These models show their chain of thought between reasoning tags before providing the final answer. This is exactly the behavior inherited from the R1 teacher model. In practice, an R1-Distill-Qwen-14B works through structured reasoning on a math problem where a classic Qwen 14B would answer more superficially — with the same memory footprint.
#The limits: what cannot be transferred
Distillation has a ceiling set by the student's size. You can't pour a giant's ocean into a thimble. Understanding what gets left behind prevents disappointment and poor model choices.
- Fine-grained factual knowledge
- The small model cannot remember everything the teacher knows. On obscure or niche facts, it is more likely to hallucinate.
- Off-topic robustness
- For tasks far removed from what was distilled, the student falls back to roughly the level of a standard model of the same size.
- The capacity ceiling
- A distilled 7B model is still a 7B model. Distillation optimizes parameter usage; it doesn’t add any.
- Consistency over a very long context
- Following a chain of reasoning across tens of thousands of tokens remains harder for a small model, distilled or not.
There is also a subtler risk: a distilled model can imitate the teacher's reasoning style without always matching its accuracy. It “acts as if it were thinking” convincingly, but may confidently get a problem beyond its ability wrong. A polished chain of thought is not a guarantee of accuracy—it is something to always verify, especially for important decisions.
#Spot a good distillation
Not all distilled models are equal. Before downloading one, a few indicators can help you seriously assess its quality without being an expert.
- An identified professor
- A good model card clearly identifies the teacher model (e.g., “distilled from R1”). A reputable teacher is a good starting sign.
- A clear foundation
- What architecture is the model built on (Qwen, Llama…)? A solid, well-supported base makes local use easier.
- Targeted benchmarks
- Look at scores on the target tasks (math, code, reasoning), not some vague overall number. And compare it with the same-size undistilled model.
- Size-to-promise consistency
- Be wary of a tiny model marketed as equivalent to a giant model at everything. Distillation helps, but it doesn't work miracles.
- The reasoning format
- For a reasoning distillate, verify that it genuinely exposes its chain of thought and that it is relevant, not merely decorative.
#The leading distillations to try locally
Here are the distilled models worth considering in 2026, from lightest to most demanding, with their indicative Q4_K_M footprint.
- DeepSeek-R1-Distill-Qwen-1.5B
- ≈ 1.5 GB. The tiny-model gamble: runs even on a CPU or small laptop. Impressive at math for its size.
- DeepSeek-R1-Distill-Qwen-7B
- ≈ 5 GB, on an 8 GB card. The ideal entry point for discovering distilled reasoning in everyday use.
- DeepSeek-R1-Distill-Qwen-14B
- ≈ 9 GB, the target for a RTX 3060 12 GB or 4070. The best accessible reasoning-to-hardware ratio.
- DeepSeek-R1-Distill-Qwen-32B
- ≈ 19 GB, for a RTX 4090 24 GB. The closest to the teacher model if your VRAM allows it.
If you are just starting, begin with the 7B or 14B version: they show the value of distillation on a single consumer card, without the bulk of large configurations. The 1.5B is an excellent sandbox for understanding reasoning behavior even without a GPU.
#Try a distilled model in 3 minutes
If Ollama is already installed and listening on its default port (http://localhost:11434), two commands are enough to get a feel for what distillation brings compared with a classic model of the same size.
- 01Download and run a distilled modelGet an R1-Distill and chat with it. The first launch downloads the model (a few GB depending on the selected size).
- 02Pose a reasoning problemGive it a multi-step math or logic problem. Observe its chain of thought: it works through its reasoning before reaching a conclusion.
- 03Compare with the non-distilled modelRun the same-size Qwen classic and ask the same question. The distilled model should reason more structurally, with the same VRAM and speed.
#Go further
Distillation is easier to understand when connected to the other fundamentals of local AI. These guides directly build on this one:
- Install DeepSeek R1 with Ollama
- The step-by-step guide to running the distilled 7B/14B/32B versions locally, with the VRAM requirements for each size.
- MoE explained: why a 30B-A3B runs like a small model
- The other major architectural trick that makes large models accessible locally, complementary to distillation.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Once you have chosen your distilled model, quantization determines its final memory footprint on your card.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.