AI System Architecture Gallery
Interactive SVG diagrams for all major AI systems — RAG pipelines, agent patterns, training architectures, and more. Drag nodes · scroll to zoom · click for flow details.
Naive RAG Pipeline
The simplest RAG architecture: chunk documents, embed them, store in a vector DB, retrieve the top-k, and pass as context to the LLM.
Advanced RAG — Hybrid Search + Re-ranking
Production RAG adds query rewriting, hybrid search (dense + sparse BM25), cross-encoder re-ranking, and HyDE for better retrieval quality.
GraphRAG — Knowledge Graph Retrieval
Microsoft GraphRAG builds a knowledge graph from documents. Community detection enables global reasoning across an entire corpus — impossible with chunk retrieval.
Agentic RAG — ReAct Loop
Agentic RAG gives the LLM autonomy to decide when to retrieve, what to query, and whether to retrieve again — using the Reason-Act-Observe loop.
Supervisor-Worker Multi-Agent Pattern
A supervisor LLM decomposes tasks and routes them to specialized worker agents. LangGraph implements this as a directed state graph with conditional edges.
LangGraph State Machine Flow
LangGraph models multi-agent systems as directed graphs. Nodes are agents/tools; edges are conditional on the LLM's decision. State persists across all nodes.
Model Context Protocol (MCP) Architecture
MCP standardizes how LLMs connect to external tools and data. Clients (Cursor, Claude) connect to MCP Servers via stdio or HTTP+SSE transport using JSON-RPC 2.0.
A2A Protocol — Agent-to-Agent Communication
Google's A2A protocol enables agents built on different frameworks to delegate tasks via HTTP. Agents advertise capabilities via AgentCards and communicate structured Tasks.
Scaled Dot-Product Self-Attention
Each token attends to all others by computing Query·Key scores, scaling by √dₖ, applying softmax to get weights, then multiplying Values. This runs h times in parallel (Multi-Head).
KV Cache — Autoregressive Decoding
During inference, keys and values for all past tokens are cached so only the new token's Q is computed each step — reducing decoding from O(n²) to O(n) per step.
LoRA — Low-Rank Adaptation Architecture
LoRA freezes the pre-trained weights W₀ and injects trainable low-rank matrices A (d×r) and B (r×d). Only A and B are updated, reducing trainable params by 10,000×.
DPO vs RLHF Post-Training Pipeline
RLHF requires training a reward model then running PPO. DPO eliminates both by directly optimizing the LLM using preference pairs — fewer moving parts, more stable.
LLaVA VLM Architecture
The standard VLM pattern: a frozen vision encoder (ViT-L) extracts patch tokens, a projection layer maps them to text embedding space, then the LLM generates conditioned on both.
ColPali — Multimodal Document RAG
ColPali embeds document page images directly using a VLM (PaliGemma), producing multi-vector representations. MaxSim scoring retrieves visually-rich documents without any OCR.
MoE Expert Routing (Mixtral / DeepSeek Pattern)
The router selects top-k experts for each token. Most parameters are inactive per token — giving N× more capacity at the same compute cost as a dense model.
Whisper ASR Pipeline
Whisper converts audio → log-Mel spectrogram → convolutional stem → encoder → autoregressive decoder with task conditioning tokens.
Knowledge Distillation — Teacher-Student Training
The student is trained on soft labels (temperature-scaled logits) from the teacher plus the hard ground-truth labels. This transfers the teacher's knowledge without its size.
Production AI Security Stack
A layered defense for LLM APIs: input guardrails → authentication → RBAC → rate limiting → PII masking → LLM call → output guardrails → audit log.
LLMOps CI/CD Pipeline
LLM CI/CD extends standard DevOps with prompt regression tests, model evaluation, cost estimation, and container builds — gating on quality before deployment to EKS.
LLM Data Flywheel — Llama 3 Pattern
Iterative self-improvement: SFT → generate responses → reward model scores → rejection sampling → better SFT data → repeat. Each iteration improves the model.