✏️
DATA ~15 min6 concepts · 5 quiz questions

Synthetic Data Engineering

Self-Instruct, Evol-Instruct, LLM-as-Judge scoring, deduplication, quality filtering pipelines.

Overview

High-quality training data is more valuable than model architecture choices. This module covers the modern synthetic data pipeline: generating diverse instruction data with LLMs (Self-Instruct, Evol-Instruct), scoring quality with LLM-as-Judge, deduplicating with MinHash LSH, and filtering with perplexity and reward models — the techniques behind Llama 3 and Mistral's data engines.

Concept Flashcards

6 cards — click to flip and test recall

1 / 6
0/6 mastered
Concept

Self-Instruct

Click to reveal explanation
Click to reveal answer
Explanation

Self-Instruct (Wang et al., 2023) bootstraps instruction data: start with 175 human-written seed tasks, use GPT to generate new (instruction, input, output) tuples, filter low-quality ones, add to the pool, and repeat. Achieved 90%+ of InstructGPT quality using only GPT-3 and 5% human-written seeds.

Click to flip back

System Architecture

1 interactive diagram — drag nodes · scroll to zoom · click for details

Iterative self-improvement: SFT → generate responses → reward model scores → rejection sampling → better SFT data → repeat. Each iteration improves the model.

⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
expandevolved promptscandidatestop-kbetter model → better dataSeed Data(human-written)Evol-Instruct(LLM generates more)LLM GeneratesResponsesReward Model(scores quality)Rejection Sampling(top 20-30%)Dedup + Filter(MinHash LSH)SFT Training(next iteration)
Instruction Generation
Reward Scoring
Rejection Sampling
Iterative Training

Key Concepts

6 concepts — click to expand

Self-Instruct (Wang et al., 2023) bootstraps instruction data: start with 175 human-written seed tasks, use GPT to generate new (instruction, input, output) tuples, filter low-quality ones, add to the pool, and repeat. Achieved 90%+ of InstructGPT quality using only GPT-3 and 5% human-written seeds.

Tech Stack

distilabelDataDreamerArgillaHuggingFace HubDoclingLlamaParse
Ready to test yourself?
5 questions · score ≥ 60% to mark complete