MIT Interactive Diagrams: From Attention to Mixtral and DeepSeek-V3 Architectures
vtabbott_ · x · 2026-10-03
A user shares an interactive Transformer architecture diagram tool from MIT (zardini.mit.edu) that lets you explore the internals of many models:
- Building blocks: scaled dot-product attention with its training step, causal self-attention with weights and a residual connection, multi-head attention, and grouped-query attention (GQA);
- Classic models: Attention Is All You Need (2017), Mixtral-8x7B sparse MoE (2023), DeepSeek-V3 latent attention + MoE (2024);
- Modern models: Kimi K3's delta attention and attention residuals, GLM-5.3's MoE + sparse attention, DeepSeek-V4.1-Flash, and MiMo-V2.6-Pro's sliding-window attention + MoE;
- Interactivity: switch between Decode and Cached modes to see where the KV-cache should go, with an expressible relationship between prefill and inference forms; toggle quantised vs unquantised (FP vs real numbers) views, visualize arrows and broadcasting, and open the accompanying notebook.
The author notes that abstract representations let you derive where the KV-cache belongs.
More from coding & agent
- Vesence launches Agent-Native Desktop; YC's Garry Tan calls it the biggest AI UI leap yet — garrytan · 2026-10-03
- Dev uses Opus 5.5 to produce Agent Lens explainer video, calls it 'ridiculously good' — _ScottCondron · 2026-10-03
- Agent Lens design: embed all agent conversations to surface top failures worth fixing — _ScottCondron · 2026-10-03
- OpenClaw adds Tencent's AI-Infra-Guard to ClawHub's skill security review pipeline — heyneighbor · 2026-10-03
- Hermes Agent adds one-click MCP install: 4 steps to wire in DeepWiki — GrowthHackingEU · 2026-10-03
- Seroter Daily Reading List: AI coding assistants reshape engineering, Spanner Omni hits GA — rseroter · 2026-10-03