training
Knowledge Distillation
A specific form of model compression where a smaller student model is trained to match the output probability distribution of a larger teacher model. Knowledge distillation transfers the teacher's dark knowledge, including information about incorrect class probabilities.
In practice
DistilGPT-2 was trained using knowledge distillation from GPT-2, achieving 95% of its performance at half the size.