Overview
High-quality training data is more valuable than model architecture choices. This module covers the modern synthetic data pipeline: generating diverse instruction data with LLMs (Self-Instruct, Evol-Instruct), scoring quality with LLM-as-Judge, deduplicating with MinHash LSH, and filtering with perplexity and reward models — the techniques behind Llama 3 and Mistral's data engines.
Concept Flashcards
6 cards — click to flip and test recall
1 / 60/6 mastered
System Architecture
1 interactive diagram — drag nodes · scroll to zoom · click for details
Iterative self-improvement: SFT → generate responses → reward model scores → rejection sampling → better SFT data → repeat. Each iteration improves the model.
⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
Instruction Generation
Reward Scoring
Rejection Sampling
Iterative Training
Key Concepts
6 concepts — click to expand
Self-Instruct (Wang et al., 2023) bootstraps instruction data: start with 175 human-written seed tasks, use GPT to generate new (instruction, input, output) tuples, filter low-quality ones, add to the pool, and repeat. Achieved 90%+ of InstructGPT quality using only GPT-3 and 5% human-written seeds.
Tech Stack
distilabelDataDreamerArgillaHuggingFace HubDoclingLlamaParse
Ready to test yourself?
5 questions · score ≥ 60% to mark complete