DeltaNet notes unpack Kimi Delta Attention with equations and a state-update diagram
nrehiew_ · x · 2026-07-22
A technical thread breaks down DeltaNet / Gated DeltaNet / Kimi Delta Attention from first principles, using equations and a diagram to show how the state update works.
- It starts from linear attention by removing softmax and rearranging matrix operations.
- It then introduces the delta rule, where the recurrent state accumulates new association information while correcting duplicated old associations.
- Gated Delta Net adds data-dependent gates (alpha, beta) to decay the previous state, erase old value associations, and write new ones.
- The image compares the update rules side by side and highlights how Kimi Delta Attention scales each dimension with a data-dependent diagonal matrix before applying the delta-style update.
- The post is framed as study notes from a conversation and points readers to the original thread for a fuller walkthrough.
More from Companies & People
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11