Overview
Pre-training gives LLMs world knowledge; fine-tuning shapes their behavior. This module covers the full post-training pipeline from supervised fine-tuning (SFT) to preference alignment (RLHF, DPO, GRPO) and parameter-efficient methods (LoRA, QLoRA, DoRA) that make it possible to adapt 70B+ models on consumer hardware.
Concept Flashcards
6 cards — click to flip and test recall
1 / 60/6 mastered
System Architecture
2 interactive diagrams — drag nodes · scroll to zoom · click for details
LoRA freezes the pre-trained weights W₀ and injects trainable low-rank matrices A (d×r) and B (r×d). Only A and B are updated, reducing trainable params by 10,000×.
⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
Frozen (pre-trained)
Matrix A (trainable)
Matrix B (trainable)
Low-rank update ΔW
Key Concepts
6 concepts — click to expand
SFT trains the model on (instruction, response) pairs using next-token prediction loss. The key challenge is data quality — models learn to mimic formatting and tone. Modern SFT datasets use chat templates (ChatML, Alpaca, Llama) to structure multi-turn conversations.
Tech Stack
HuggingFace TRLPEFTUnslothAxolotlLLaMA-FactorySageMakerTransformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete