🔧
LLM TRAINING ~15 min6 concepts · 5 quiz questions

Advanced Fine-Tuning

LoRA, QLoRA, DoRA, SFT, DPO, GRPO, ORPO — the complete post-training pipeline.

Overview

Pre-training gives LLMs world knowledge; fine-tuning shapes their behavior. This module covers the full post-training pipeline from supervised fine-tuning (SFT) to preference alignment (RLHF, DPO, GRPO) and parameter-efficient methods (LoRA, QLoRA, DoRA) that make it possible to adapt 70B+ models on consumer hardware.

Concept Flashcards

6 cards — click to flip and test recall

1 / 6
0/6 mastered
Concept

Supervised Fine-Tuning (SFT)

Click to reveal explanation
Click to reveal answer
Explanation

SFT trains the model on (instruction, response) pairs using next-token prediction loss. The key challenge is data quality — models learn to mimic formatting and tone. Modern SFT datasets use chat templates (ChatML, Alpaca, Llama) to structure multi-turn conversations.

Click to flip back

System Architecture

2 interactive diagrams — drag nodes · scroll to zoom · click for details

LoRA freezes the pre-trained weights W₀ and injects trainable low-rank matrices A (d×r) and B (r×d). Only A and B are updated, reducing trainable params by 10,000×.

⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
forwardforwardW₀xbaseα/r · ΔWInput xW₀ (frozen)d × dA (trainable)d × rB (trainable)r × dW₀xΔW = BA·x(low-rank update)+ (merge)Output h
Frozen (pre-trained)
Matrix A (trainable)
Matrix B (trainable)
Low-rank update ΔW

Key Concepts

6 concepts — click to expand

SFT trains the model on (instruction, response) pairs using next-token prediction loss. The key challenge is data quality — models learn to mimic formatting and tone. Modern SFT datasets use chat templates (ChatML, Alpaca, Llama) to structure multi-turn conversations.

Tech Stack

HuggingFace TRLPEFTUnslothAxolotlLLaMA-FactorySageMakerTransformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete