Interactive Attention diagrams page gets major update, from scaled dot-product to DeepSeek latent attention
vtabbott_ · x · 2026-09-28
vtabbott, teaming up with MIT's Gioele Zardini, has rolled out a major update to the interactive Transformer diagram site at zardini.mit.edu, breaking down attention mechanics through interactive arrows-and-boxes diagrams.
Highlights:
- Attention fundamentals: scaled dot-product attention with its training step, causal self-attention with projections and residual connections, multi-head attention, and grouped-query attention (GQA)
- Classic models: the original Transformer (Attention Is All You Need), Mixtral-8x7B sparse MoE, and DeepSeek-V3 with latent attention + MoE
- Modern models: diagrams covering GLM and DeepSeek-V4.1-Flash architectures
- Diagrams support Forward/Decode/Cached pass views and distinguish real-number math from the quantized formats in released code
A standout new feature: side-by-side comparison of standard graphs versus broadcasted diagrams, where "the wires represent axes," building intuition from simple arrows-and-boxes up to fully broadcasted forms. A series of explainer posts is planned.
More from Research
- Neuroscientist pushes back: task-trained nets converge on brain-like representations that can predict primate cortex — coherence · 2026-09-28
- ESPnet's YODAS3 Speech Dataset with 1M+ Entries Trends on Hugging Face — espnet · 2026-09-28
- BoundInk Treats Inter-Character Boundaries as Explicit Units, Cutting DTW by up to 47.8% — SUNGKYUNARCH · 2026-09-28
- TUM's Continuous Depth Batching Unlocks 99% of Speedup for Looped LMs — TUM · 2026-09-28
- ARGUS: An LLM Pipeline Audits Causal Identification Assumptions, Catching 73% of Planted Flaws — Yonghong Zhang · 2026-09-28
- Category theory meets deep learning: PyNCD diagrams derive hardware-aware FlashAttention — GioeleZardini · 2026-09-28