Prof. Tom Yeh Publishes Kimi 3 Seminar Tracing Attention from Vanilla to Kimi Delta Attention
ProfTomYeh · x · 2026-09-04
Prof. Tom Yeh has released the recording of his Kimi 3 seminar, working through the evolution of attention mechanisms by hand in Excel: vanilla attention → local attention → linear attention → delta attention → Kimi Delta Attention (KDA).
Nathan Lambert, author of Interconnects and the RLHF book, joined as a special guest and shared behind-the-scenes stories from his visit to Moonshot AI. The session also covers Kimi 3's architecture, open vs closed models, and audience Q&A.
More from Models
- Respan launches P-1 privacy model, beats AWS on PII detection F1 across benchmarks — CodeByPoonam · 2026-09-04
- OpenAI repo claims GPT-6-Astra Lean-formalized a prime-gaps bound of 186 — scaling01 · 2026-09-04
- Traces of gpt-6-astra spotted in Codex, fueling next-model rumors — drinksbeerdaily · 2026-09-04
- Gemini, Grok and Claude went down together — and it wasn't OpenAI's fault — iiiaaa2022 · 2026-09-04
- GPT-6-Astra Rumored to Land Today, With Another Big Jump From Bel Later This Year — koltregaskes · 2026-09-04
- Meta's Muse Spark 1.3 tops DeepSWE coding benchmark at 1/5 the price of Opus 5 — davidthesong · 2026-09-04