📓
COMPRESSION ~15 min6 concepts · 5 quiz questions

Knowledge Distillation

Student-Teacher paradigm, KL Divergence & Attention Transfer losses, GGUF quantization for edge deployment.

Overview

Knowledge distillation transfers learned representations from a large teacher model to a smaller student model. Combined with quantization techniques like GGUF and GPTQ, distillation enables powerful models to run on edge devices and consumer hardware. This module covers the full model compression pipeline.

Concept Flashcards

6 cards — click to flip and test recall

1 / 6
0/6 mastered
Concept

Student-Teacher Framework

Click to reveal explanation
Click to reveal answer
Explanation

The teacher is a large, well-trained model. The student is a smaller architecture trained to match the teacher's output distribution (soft labels). Soft labels contain rich inter-class relationships not present in hard one-hot labels, making them a powerful training signal.

Click to flip back

System Architecture

1 interactive diagram — drag nodes · scroll to zoom · click for details

The student is trained on soft labels (temperature-scaled logits) from the teacher plus the hard ground-truth labels. This transfers the teacher's knowledge without its size.

⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
forward passforward passstudent dist.α ×backpropTraining DataTeacher Model(large, frozen)Student Model(small, trained)Soft Logits(Temp T=4)Student LogitsGround Truth(one-hot)KL DivergenceLossCross-EntropyLossTotal Lossα·KL + (1-α)·CE

Key Concepts

6 concepts — click to expand

The teacher is a large, well-trained model. The student is a smaller architecture trained to match the teacher's output distribution (soft labels). Soft labels contain rich inter-class relationships not present in hard one-hot labels, making them a powerful training signal.

Tech Stack

PyTorchllama.cppOllamavLLMLiteLLMHuggingFace Transformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete