Cookbook · FineTuning
Distillation
5 mindistillationfine-tuningtrainingkl-divergence
The idea, in one analogy
If you wanted to teach a student a subject, you could just give them the textbook's final answers ("the answer to problem 4 is 12"). Or you could show them the expert's full reasoning and confidence: "12 is almost certainly right, though 11 was a plausible near-miss and 13 is clearly wrong." The second version carries far more information per example. Knowledge distillation is training a small "student" model on a large "teacher" model's full output distribution, rather than just the single correct answer.
Hard labels vs. soft labels
Hard-label distillation: the student is trained on the teacher's single best output (e.g., the teacher generates an answer, and the student trains on that generated text as if it were ground truth). Simple, and works with any teacher you can only query as a black box.
Soft-label distillation: the student is trained to match the teacher's full probability distribution over possible next tokens, not just its top pick, usually via KL divergence between the two distributions. This carries much more signal (the relative confidence across all the runner-up tokens), but requires access to the teacher's actual output logits, not just its generated text, which rules out any teacher you can only access through a text-only API.
A worked example: what "dark knowledge" looks like
For the prompt "the cat sat on the ___", a teacher model's hard label is just the single most likely word. But its full soft-label distribution over candidate words carries much more:
| Candidate | Teacher's probability |
|---|---|
| mat | 52% |
| rug | 23% |
| floor | 14% |
| chair | 6% |
| airplane | ~0% |
The hard label alone throws away everything except "mat." The soft label tells the student that "rug" and "floor" are also reasonable in this context, "chair" is a stretch but not absurd, and "airplane" is essentially ruled out, information a single correct-answer label can never carry. This relative ranking across the wrong answers is what Hinton et al. called "dark knowledge": most of what the teacher has learned about a concept is encoded in how it distributes probability across the incorrect options, not just in which option wins.
The constraint that gets skipped: vocabulary compatibility
Soft-label distillation requires the teacher and student to share the same tokenizer, or at least a way to align token boundaries. Otherwise the teacher's probability distribution is over a different vocabulary than the student's output layer, and there's no direct way to compute a KL divergence between them. If your student uses its own custom tokenizer (see tokenization) and the teacher uses a different one, soft-label distillation is off the table until that mismatch is resolved, either by using the teacher's tokenizer for the student too, or by a token-alignment scheme. Verify this before designing a distillation pipeline around it, rather than discovering it mid-project.
The other cost: generating teacher data isn't free
Whether you're doing hard-label or soft-label distillation, you generally need the teacher to run inference to produce training signal, and autoregressive generation is decode-bound (see inference for why), much slower than a single forward pass over an existing sequence. On modest hardware, generating enough teacher outputs to meaningfully train a student can be the bottleneck of the whole project, more so than the student's training step itself. Measure teacher throughput directly before committing to a distillation plan's scope.
Where to look further
- Hinton et al., "Distilling the Knowledge in a Neural Network": the original distillation paper, including temperature scaling.
- Hugging Face's distillation examples: a maintained reference implementation.
- This project's own distillation investigation, including the vocabulary-mismatch finding above:
slm-from-scratchPhase 4.