1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Knowledge Distillation: Large to Small

Train a small, fast model to mimic a large teacher — economics, pipeline, and quality filters for production distillation.

Last updated

After this section you can

  • Distinguish logit distillation from sequence-level distillation and say which one an API allows
  • Build a teacher-to-student pipeline with a filter independent of the teacher
  • Decide when distillation beats simply prompting a smaller off-the-shelf model
49

Knowledge Distillation: Large to Small

A large teacher model generates the training data; a small student learns to reproduce it on one task. You keep much of the quality at a fraction of the serving cost, if the filter is strict and the baseline is honest.

Key idea

Distillation is fine-tuning where a stronger model writes the labels. It pays when a small off-the-shelf model falls well short on a narrow, high-volume task and a large model does it well. The student can only be as good as the filtered teacher data.

A large teacher labels real prompts; a filtered set trains a small student
SEQUENCE-LEVEL DISTILLATION: THE TEACHER WRITES THE TRAINING SET 1 · Real prompts sampled from traffic, hard tail included 2 · Teacher large model answers each one 3 · Filter verify, judge, dedup drops a real share 4 · Dataset prompt → answer pairs, versioned 5 · Student small model, LoRA fine-tune 6 · Evaluate vs teacher held-out real prompts · cost · latency below target: more or better teacher data Whatever the filter lets through becomes the student’s behaviour, mistakes included. Stage 3 decides quality.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium