Prof. Tom Yeh Posts Kimi 3 Seminar Recording with Nathan Lambert on KDA Architecture
ProfTomYeh · x · 2026-09-21
Prof. Tom Yeh uploaded the full recording of his Kimi 3 Frontier AI Seminar, continuing his 'AI by Hand' Excel-based series.
Key points:
- Guest Nathan Lambert (Interconnects author, RLHF book) shared behind-the-scenes stories from his visit to Moonshot AI.
- Main storyline traces the Linear Transformer lineage from Linear Attention and DeltaNet to Kimi Delta Attention (KDA) in Kimi 3.
- Agenda included open vs closed models, audience Q&A, and hand-derivation of attention mechanics.
- The seminar series previously covered DeepSeek mHC, Qwen 3.6, Gemma 4, and Google Ironwood TPU.
More from Models
- MiMo near-SOTA on DeepSWE with just ~$2.6M RL run: will data cost more than training? — my_cat_can_code · 2026-09-21
- humansand's Persimmon model learns to share info gradually like humans, with Trickle Test — niloofar_mire · 2026-09-21
- Why yes/no answers are fast for LLMs: output tokens dominate latency — tinyfool · 2026-09-21
- ChatGPT reportedly removes free-tier chat limits, offering unlimited text chats — nikola_mr64990 · 2026-09-21
- Users say top-tier Astra is too costly, hope GPT-6 fixes token economics — CtrlAltDwayne · 2026-09-21
- TypeSafe's JEV model fully open with $5 free credit, powers 1-second 3D scene generation — tinyfool · 2026-09-21