# How a small model learns from a big one > Many of the small open models you can run at home were taught by a bigger model, not just by the internet. What distillation is, why it works so well, and what it can't hand down. Source: https://rfriedmann.de/blog/how-small-models-learn-from-big-ones/ Published: 2026-10-10 · Track: learn · Level: Beginner Read the release notes of a small open model and a word keeps turning up: *distilled*. A 1B or 3B model that holds a conversation far better than its size suggests, and somewhere in the small print it says it was trained with help from a much larger sibling. That help is not a figure of speech. The big model was used as a teacher, and the small one learned from it directly, in a way that turns out to be far more efficient than learning from raw text alone. It is one of the main reasons the models that fit [on a single card](/blog/how-a-big-model-fits-on-one-card/) got as good as they are. ## What ordinary training throws away Start with how a model normally [learns](/blog/how-a-model-learns/). It reads a piece of text, guesses the next word, gets told which word actually came next, and nudges its weights toward that answer. Repeat a few trillion times. The feedback in that loop is thin. For "The weather today is ___", the training text says "sunny", so "sunny" is right and every other word in the vocabulary is equally wrong: "cold" counts exactly as much against it as "banana". The model has to work out on its own, over a vast number of examples, that some wrong answers are much more reasonable than others. A trained model already knows this. Ask it about the same sentence and you do not get one word, you get a [whole spread of scored options](/blog/why-different-answers/): "sunny" likely, "cold" and "lovely" quite plausible, "grey" possible, "purple" essentially never. That spread is information the original text never contained.
"The weather today is ___" - two kinds of answer key
[diagram omitted — see the page for the chart]
The same training example, graded two ways. The text only says which word appeared. The teacher says how close every other word came, and that "cold" is a far better miss than "grey". A student learning from the bottom row gets several lessons out of every sentence. Numbers illustrative.
## Learning from the teacher's doubts Distillation uses that spread as the answer key. The student reads the same text, makes its own guess, and is corrected not toward "sunny, full stop" but toward the teacher's full distribution: mostly sunny, a fair chance of cold, a little lovely. The idea was put plainly in a 2015 paper by Geoffrey Hinton and colleagues, and Hinton gave the extra information a memorable name: *dark knowledge*. It is everything the teacher knows that never shows up in its top answer. That "cold" is a sensible word here. That a picture of a dog looks a bit like a cat and nothing like a car. The relationships between wrong answers are where much of the understanding lives, and ordinary training only ever reveals them indirectly. Two practical consequences follow: - **Fewer examples go further.** Each sentence now carries a graded lesson about thousands of possible words instead of a single right-or-wrong. The student extracts more from the same text. - **The student learns what the teacher learned, not what the text says.** If the teacher has absorbed that a sentence is ambiguous, the student is taught the ambiguity directly instead of having to rediscover it. There is a tuning detail in the original recipe that rhymes with something you may already know. The teacher's spread is softened with a [temperature](/blog/why-different-answers/) setting before the student sees it, so the small probabilities of the unlikely-but-sensible words are big enough to learn from. Same dial as in sampling, used for the opposite purpose: here the point is to make the long shots visible. ## Two ways to do it today The textbook version above needs the teacher's full scores for every word, at every position. That is the strong form, and it requires having the teacher's weights or at least its raw output. The other form only needs the teacher's text.

Match the scores

Learn from its writing

Both are in heavy use among the models you can download: - **Gemma 2** trained its 2B and 9B models by distillation from a larger model rather than plain next-word prediction, and Google said so prominently. - **Llama 3.2 1B and 3B** were pruned down from Llama 3.1 8B and then pretrained with the output scores of the 8B and 70B Llama 3.1 models as targets. - **DeepSeek-R1's small variants**, from 1.5B to 70B, are ordinary Qwen and Llama models fine-tuned on around 800,000 examples curated with R1, most of them worked solutions from its training run. No reinforcement learning of their own: they learned to [think before answering](/blog/why-models-think-before-answering/) by reading a big model doing it. - **DistilBERT**, an older and much-cited case, came out 40% smaller and 60% faster than the BERT model it was distilled from while keeping, by its authors' measure on the GLUE benchmark, 97% of its language understanding. That last pattern, roughly most of the ability for a fraction of the size, is the whole commercial point. Every token you generate pays for every weight in the model, so a student that is a tenth the size and nearly as good is a tenth the cost to run, every day, forever. Training it once is cheap by comparison. ## What doesn't come along The "nearly as good" deserves a closer look, because the gap is not spread evenly. **Facts compress worst.** A model's knowledge of the world is stored in its weights, and a smaller model has fewer of them. Style, format, tone and the general shape of a good answer transfer very well. Long-tail facts, the population of a small town or the signature of an obscure API, need room, and the student simply has less. Ask a distilled model something obscure and you are more likely to get a [confident invention](/blog/why-models-make-things-up/) than from its teacher. **The benchmarks the teacher was tuned for transfer best.** If the distillation data was heavy on maths and code, the student will look great on maths and code. That says much less about your use, which is why [testing on your own work](/blog/testing-a-model-on-your-own-work/) matters more for small models, not less. **The teacher's flaws come too.** A student cannot learn to be more correct than the answers it imitates. In the text-only form especially, the teacher's errors arrive looking exactly like its successes. **Not everyone may be used as a teacher.** Training on another model's output is technically easy and contractually messy: several commercial providers forbid using their outputs to build competing models in their terms of service. Open-weights teachers with permissive licences are the clean route, and that is a large part of why model families now ship in several sizes at once. The big one is the product and the teacher. ## Why this matters if you run models yourself When you read the [spec sheet](/blog/reading-model-specs/) of a small model, the parameter count tells you its running cost, not where its ability came from. A 3B model distilled from a strong teacher can be far more useful than a 7B model trained on text alone, and the size in the name gives you no hint. So treat "distilled" as a good sign with a known shape: fluent, well-formatted, often surprisingly capable at structured tasks, thinner on facts than the teacher. Pair it with [your own documents](/blog/your-own-documents/) when facts matter, check it on your own work, and the small model on your own hardware does a lot of what the large one does, because in a very literal sense, it was taught by it.