Knowledge Distillation: Large to Small
Train a small, fast model to mimic a large teacher — economics, pipeline, and quality filters for production distillation.
Last updated
After this section you can
- Distinguish logit distillation from sequence-level distillation and say which one an API allows
- Build a teacher-to-student pipeline with a filter independent of the teacher
- Decide when distillation beats simply prompting a smaller off-the-shelf model
Knowledge Distillation: Large to Small
A large teacher model generates the training data; a small student learns to reproduce it on one task. You keep much of the quality at a fraction of the serving cost, if the filter is strict and the baseline is honest.
Distillation is fine-tuning where a stronger model writes the labels. It pays when a small off-the-shelf model falls well short on a narrow, high-volume task and a large model does it well. The student can only be as good as the filtered teacher data.