Overview
Mixture of Experts (MoE) models like Mixtral, DeepSeek-V3, and GPT-4 (rumored) achieve dense-model performance while activating only a fraction of parameters per token. This module covers the gating mechanism, expert routing, load balancing challenges, and the fundamental tradeoffs that make MoE models both powerful and complex to deploy.
Concept Flashcards
6 cards — click to flip and test recall
1 / 60/6 mastered
System Architecture
1 interactive diagram — drag nodes · scroll to zoom · click for details
The router selects top-k experts for each token. Most parameters are inactive per token — giving N× more capacity at the same compute cost as a dense model.
⊕Scroll to zoom · Drag nodes · Drag canvas to pan
100%
Active experts (top-k=2)
Inactive experts
Router (gating)
Weighted merge
Key Concepts
6 concepts — click to expand
In a standard transformer, each token passes through a single FFN. In MoE, the FFN is replaced by N expert FFNs, and a router selects top-k (usually k=2) experts per token. Only k/N fraction of parameters are active, but total parameter count is N× larger — enabling massive scale at fixed compute.
Tech Stack
PyTorchvLLMSGLangHuggingFace TransformersOllama
Ready to test yourself?
5 questions · score ≥ 60% to mark complete