DeltaNet notes unpack Kimi Delta Attention with equations and a state-update diagram
nrehiew_ · x · 2026-07-22
A technical thread breaks down DeltaNet / Gated DeltaNet / Kimi Delta Attention from first principles, using equations and a diagram to show how the state update works.
- It starts from linear attention by removing softmax and rearranging matrix operations.
- It then introduces the delta rule, where the recurrent state accumulates new association information while correcting duplicated old associations.
- Gated Delta Net adds data-dependent gates (alpha, beta) to decay the previous state, erase old value associations, and write new ones.
- The image compares the update rules side by side and highlights how Kimi Delta Attention scales each dimension with a data-dependent diagonal matrix before applying the delta-style update.
- The post is framed as study notes from a conversation and points readers to the original thread for a fuller walkthrough.
More from Companies & People
- Jensen Huang Defends Kimi K3: Claims Everyone Got the Logic Backwards — pstAsiatech · 2026-07-22
- Where Did All the Computer-Science Professors Go? — ArtificialOther · 2026-07-22
- Every says a $625-a-year membership added $9,000 in MRR in two days — danshipper · 2026-07-22
- Scott Belsky says vertical workflow interfaces are the real moat above interchangeable models — round · 2026-07-22
- The New Role of Engineers in the AI Era: Shaping Marketing and Brand Narrative — emilyzsh · 2026-07-22
- ChatGPT’s market share falls below 50% for the first time, TechCrunch reports — KeanuRave100 · 2026-07-22