Overview
Knowledge distillation transfers learned representations from a large teacher model to a smaller student model. Combined with quantization techniques like GGUF and GPTQ, distillation enables powerful models to run on edge devices and consumer hardware. This module covers the full model compression pipeline.
Concept Flashcards
6 cards — click to flip and test recall
1 / 60/6 mastered
System Architecture
1 interactive diagram — drag nodes · scroll to zoom · click for details
The student is trained on soft labels (temperature-scaled logits) from the teacher plus the hard ground-truth labels. This transfers the teacher's knowledge without its size.
⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
Key Concepts
6 concepts — click to expand
The teacher is a large, well-trained model. The student is a smaller architecture trained to match the teacher's output distribution (soft labels). Soft labels contain rich inter-class relationships not present in hard one-hot labels, making them a powerful training signal.
Tech Stack
PyTorchllama.cppOllamavLLMLiteLLMHuggingFace Transformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete