Beginner 8 minConcepts

Distillation: how small models inherit gros

You’ve probably come across names like DeepSeek-R1-Distill-Qwen-14B and wondered why a “small” model suddenly reasons almost as well as a giant. The answer comes down to one word: distillation. It’s the technique that lets a compact model inherit the behavior of a much larger model without matching its size or slowness. This guide explains LLM distillation using the teacher-and-student analogy, details what is actually transferred and what gets lost along the way, and shows how to recognize successful distillation before running it on your own system.

By Clara M.·Update 2026-08-07·Tested on Windows, macOS, and Linux

#The problem distillation solves

The best open models are huge: several hundred billion parameters. They're brilliant, but unusable on a consumer machine: they require tens of gigabytes of VRAM and generate tokens slowly. By contrast, a 7B or 14B model fits on a gaming card and responds quickly, but it often lacks the larger model's reasoning finesse.

So we would like to recover some of a large model’s abilities in a format that runs locally. The naive approach would be to simply retrain a small model on the same raw data as the large one. That is not enough: at a smaller scale, learning alone from an ocean of text produces a decent model, but not an exceptional one. Distillation offers a shortcut: instead of relearning the world from scratch, the small model learns directly from the large one.

i
The intuition in one sentence
Distillation is not about compressing a large model. It is about training a small model to imitate a large model's answers and reasoning style—like an apprentice watching the master work.

#The teacher-student principle

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Distillation involves two models. The “teacher” is the highly capable large model we want to imitate. The “student” is the small model we train. The principle is simple: ask the teacher thousands of questions, collect its answers, and train the student to produce the same answers. The student no longer learns from raw text found on the web, but from examples already “chewed up” by an expert model.

The subtlety that makes all the difference is exactly what the student copies. In the richest version, they do not imitate only the final word chosen by the teacher, but the entire probability distribution for the next word: the teacher does not merely say “the answer is cat,” but indicates “cat very likely, dog somewhat likely, toaster virtually impossible.” These nuances—sometimes called soft labels—contain information that the simple correct answer does not provide.

Teacher (teacher)
The large reference model, expensive but excellent. It serves as the source; it is not deployed on your system.
Student (student)
The small model trained to reproduce the professor’s behavior. This is the one you will run locally.
Soft labels
The teacher's nuanced probabilities for every possible word, not just the winning word. The richest raw material for distillation.
Distillation dataset
The full set of questions asked and the teacher's answers, which the student practices on.
→
A mental image
A student can learn independently from books, or follow the country’s best teacher, who shows them not only the right answer but also the reasoning, hesitations, and options the teacher rules out. The latter progresses much faster with the same effort. Distillation means giving that teacher to the smaller model.

#What is actually transmitted

People often imagine that distillation “transfers the knowledge” of the larger model. That's partly true, but imprecise. What transfers best is mainly behavior and ways of working, rather than raw facts. A small model has limited storage capacity: it cannot retain everything from a giant's internal encyclopedia. However, it can capture methods very well.

Reasoning style
How to break a problem into steps and formulate a chain of thought before reaching a conclusion. This is what transfers best, and it is the core of R1-Distill.
Response format
The structure, tone, and way of organizing an explanation or code. The student adopts the teacher's writing habits.
Resolution schemes
In math, code, or logic, the student inherits proven techniques: how to approach a type of problem and which checks to perform.
Part of the knowledge
Facts and associations, yes, but limited by the small model's capacity. This is the part it handles least well.

This distinction is essential for understanding recent distilled models. When DeepSeek distills R1 into a Qwen 14B, the goal is not to fit all of R1's general knowledge into 14 billion parameters—that is impossible. The goal is to transfer its way of reasoning: thinking step by step, correcting itself, and working through a proof. And that is something a small model can learn remarkably well.

i
Behavior vs. knowledge
Remember this rule: distillation mainly transfers skills and habits (how to think), not raw facts (what to know). A distilled model can reason like a giant while knowing fewer things than it does.

#Why a distilled 14B beats a standard 14B

Here is the counterintuitive part. Two models are exactly the same size—14 billion parameters—so they use the same VRAM and run at the same speed. Yet the distilled model often crushes the classic one at reasoning. The size has not changed; only the training method has. How can the same number of parameters produce a much better model?

Because training doesn't fill parameters randomly: it organizes them. A standard 14B learns on its own from raw text; it develops solid but general-purpose capabilities. A distilled 14B learns from examples produced by an elite teacher, often with complete reasoning traces. Its 14 billion parameters are wired around good methods rather than a little of everything. At equal capacity, the quality of the training signal makes the difference.

Same cost, better signal
The distilled model is neither larger nor slower, but it was trained on much higher-quality material: an expert’s outputs.
Some dense examples
The professor's answers are full of correct and complete reasoning, rare in the web's raw text.
Less noise
The web contains plenty of errors and clutter. A teacher filters it: the student learns from clean material.

Be careful not to overinterpret this: “beating” applies mainly to the tasks targeted by distillation, typically reasoning, math, and code. On pure general knowledge or fine-grained factual knowledge, the gap narrows or even reverses. A distilled model is not magically superior at everything—it is superior where it was trained to be.

#The R1-Distill case

The example that popularized local distillation is the DeepSeek-R1-Distill family released in early 2025. DeepSeek used its large R1 reasoning model as the teacher and distilled its behavior into several students of different sizes, based on existing architectures such as Qwen and Llama. The result: 1.5B, 7B, 8B, 14B, and 32B models that reason with a visible chain of thought, something previously reserved for giants.

R1-Distill-Qwen-1.5B
Tiny, runs virtually anywhere, even without a GPU. Surprisingly capable for its size on simple math.
R1-Distill-Qwen-7B / Llama-8B
The mainstream compromise: ~5 GB in Q4, on an 8 GB card. Good reasoning for everyday use.
R1-Distill-Qwen-14B
≈ 9 GB in Q4, the target for a RTX 3060 12 GB or 4070 12 GB. The sweet spot for reasoning and accessible hardware.
R1-Distill-Qwen-32B
≈ 19 GB in Q4, for a RTX 4090 24 GB. The closest to the professor, if your VRAM can handle it.

These models show their chain of thought between reasoning tags before providing the final answer. This is exactly the behavior inherited from the R1 teacher model. In practice, an R1-Distill-Qwen-14B works through structured reasoning on a math problem where a classic Qwen 14B would answer more superficially — with the same memory footprint.

!
Distilled does not mean “the real R1”
A DeepSeek-R1-Distill-Qwen-14B is not a reduced R1: it is a Qwen 14B model that learned to reason like R1. It has the method, but not all of the power or knowledge. Do not expect the performance of the full 671B model.

#The limits: what cannot be transferred

Distillation has a ceiling set by the student's size. You can't pour a giant's ocean into a thimble. Understanding what gets left behind prevents disappointment and poor model choices.

Fine-grained factual knowledge
The small model cannot remember everything the teacher knows. On obscure or niche facts, it is more likely to hallucinate.
Off-topic robustness
For tasks far removed from what was distilled, the student falls back to roughly the level of a standard model of the same size.
The capacity ceiling
A distilled 7B model is still a 7B model. Distillation optimizes parameter usage; it doesn’t add any.
Consistency over a very long context
Following a chain of reasoning across tens of thousands of tokens remains harder for a small model, distilled or not.

There is also a subtler risk: a distilled model can imitate the teacher's reasoning style without always matching its accuracy. It “acts as if it were thinking” convincingly, but may confidently get a problem beyond its ability wrong. A polished chain of thought is not a guarantee of accuracy—it is something to always verify, especially for important decisions.

i
The right reflex
Use a distilled model for what it does well — reasoning, structuring, and coding — while maintaining a critical eye toward the facts it presents. For factual reliability, RAG grounded in your documents is better than relying on the model's memory alone.

#Spot a good distillation

Not all distilled models are equal. Before downloading one, a few indicators can help you seriously assess its quality without being an expert.

An identified professor
A good model card clearly identifies the teacher model (e.g., “distilled from R1”). A reputable teacher is a good starting sign.
A clear foundation
What architecture is the model built on (Qwen, Llama…)? A solid, well-supported base makes local use easier.
Targeted benchmarks
Look at scores on the target tasks (math, code, reasoning), not some vague overall number. And compare it with the same-size undistilled model.
Size-to-promise consistency
Be wary of a tiny model marketed as equivalent to a giant model at everything. Distillation helps, but it doesn't work miracles.
The reasoning format
For a reasoning distillate, verify that it genuinely exposes its chain of thought and that it is relevant, not merely decorative.
→
The test that tells you the truth
Download the distilled model and the same-size standard model, then ask them your own difficult questions side by side. If the distilled model reasons better on your real-world cases, it's right for you — far more meaningful than a public benchmark.

#The leading distillations to try locally

Here are the distilled models worth considering in 2026, from lightest to most demanding, with their indicative Q4_K_M footprint.

DeepSeek-R1-Distill-Qwen-1.5B
≈ 1.5 GB. The tiny-model gamble: runs even on a CPU or small laptop. Impressive at math for its size.
DeepSeek-R1-Distill-Qwen-7B
≈ 5 GB, on an 8 GB card. The ideal entry point for discovering distilled reasoning in everyday use.
DeepSeek-R1-Distill-Qwen-14B
≈ 9 GB, the target for a RTX 3060 12 GB or 4070. The best accessible reasoning-to-hardware ratio.
DeepSeek-R1-Distill-Qwen-32B
≈ 19 GB, for a RTX 4090 24 GB. The closest to the teacher model if your VRAM allows it.

If you are just starting, begin with the 7B or 14B version: they show the value of distillation on a single consumer card, without the bulk of large configurations. The 1.5B is an excellent sandbox for understanding reasoning behavior even without a GPU.

#Try a distilled model in 3 minutes

If Ollama is already installed and listening on its default port (http://localhost:11434), two commands are enough to get a feel for what distillation brings compared with a classic model of the same size.

  1. 01
    Download and run a distilled model
    Get an R1-Distill and chat with it. The first launch downloads the model (a few GB depending on the selected size).
  2. 02
    Pose a reasoning problem
    Give it a multi-step math or logic problem. Observe its chain of thought: it works through its reasoning before reaching a conclusion.
  3. 03
    Compare with the non-distilled model
    Run the same-size Qwen classic and ask the same question. The distilled model should reason more structurally, with the same VRAM and speed.
Terminal
# Lancer un modèle distillé de raisonnement
ollama run deepseek-r1:14b

# Comparer avec le modèle classique de même taille
ollama run qwen3:14b
i
View the reasoning
On an R1-Distill, the response begins with a reasoning phase (chain of thought) before the conclusion. This is precisely the behavior inherited from the R1 teacher—the visible sign that distillation did its job.

#Go further

Distillation is easier to understand when connected to the other fundamentals of local AI. These guides directly build on this one:

Install DeepSeek R1 with Ollama
The step-by-step guide to running the distilled 7B/14B/32B versions locally, with the VRAM requirements for each size.
MoE explained: why a 30B-A3B runs like a small model
The other major architectural trick that makes large models accessible locally, complementary to distillation.
Choose your quantization (Q4, Q5, Q8, FP16)
Once you have chosen your distilled model, quantization determines its final memory footprint on your card.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.