Overview
Transformers are the backbone of modern AI. Understanding their internal mechanics — how attention is computed, how memory is managed, how position is encoded — is fundamental to working with any LLM. This module takes you from the basic self-attention formula all the way to modern efficiency techniques used in models like LLaMA 3 and Gemini.
Concept Flashcards
6 cards — click to flip and test recall
1 / 60/6 mastered
System Architecture
2 interactive diagrams — drag nodes · scroll to zoom · click for details
Each token attends to all others by computing Query·Key scores, scaling by √dₖ, applying softmax to get weights, then multiplying Values. This runs h times in parallel (Multi-Head).
⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
Query (Q)
Key (K)
Value (V)
Score Matrix
Softmax
Key Concepts
6 concepts — click to expand
Attention computes a weighted sum of values using query-key similarity scores. Q·Kᵀ / √dₖ prevents vanishing gradients in deep dot-products. Softmax turns scores into probabilities before weighting V.
Tech Stack
PyTorchFlash Attention 2xFormersTritonHuggingFace Transformers
Ready to test yourself?
5 questions · score ≥ 60% to mark complete