🧠
ARCHITECTURE ~15 min6 concepts · 5 quiz questions

Transformer Internals

KV Cache, Flash Attention, MHA/MQA/GQA/MLA, RoPE, Scaling Laws from first principles.

Overview

Transformers are the backbone of modern AI. Understanding their internal mechanics — how attention is computed, how memory is managed, how position is encoded — is fundamental to working with any LLM. This module takes you from the basic self-attention formula all the way to modern efficiency techniques used in models like LLaMA 3 and Gemini.

Concept Flashcards

6 cards — click to flip and test recall

1 / 6
0/6 mastered
Concept

Self-Attention & Scaled Dot-Product

Click to reveal explanation
Click to reveal answer
Explanation

Attention computes a weighted sum of values using query-key similarity scores. Q·Kᵀ / √dₖ prevents vanishing gradients in deep dot-products. Softmax turns scores into probabilities before weighting V.

Click to flip back

System Architecture

2 interactive diagrams — drag nodes · scroll to zoom · click for details

Each token attends to all others by computing Query·Key scores, scaling by √dₖ, applying softmax to get weights, then multiplying Values. This runs h times in parallel (Multi-Head).

⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
QKweightsVInput Embeddings(+ Position)Wq (Query)Wk (Key)Wv (Value)Q·Kᵀ / √dₖCausal MaskSoftmaxAttn × VWo (Output)Attention Output
Query (Q)
Key (K)
Value (V)
Score Matrix
Softmax

Key Concepts

6 concepts — click to expand

Attention computes a weighted sum of values using query-key similarity scores. Q·Kᵀ / √dₖ prevents vanishing gradients in deep dot-products. Softmax turns scores into probabilities before weighting V.

Tech Stack

PyTorchFlash Attention 2xFormersTritonHuggingFace Transformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete