Kimi K2 Architecture: Mixing KDA and Gated MLA, and the Forget-Gate Mechanism
suchenzang · x · 2026-07-28
The author analyzes Kimi K2's attention scaling, noting it mixes local KDA (Kimi Delta Attn) with global Gated MLA (Deepseek V2) in a 3:1 ratio. KDA is constrained with qk-norm, and channel-wise decay is framed as a 'forget-gate,' acting similarly to weight decay to constrain dynamic range.
More from Research
- Burkov skew AI hype: 'deterministic LLMs' and 'first agents' are old tricks rebranded — burkov · 2026-09-23
- Continuous diffusion beats discrete on random k-SAT, proposed as standard benchmark — ArashVahdat · 2026-09-23
- Grady Booch: Contemporary AI Still Lacks Abductive Reasoning, Just 'Next-Token Prediction' — Grady_Booch · 2026-09-23
- AI solves Navier-Stokes-related problem as machines upend mathematics, New Scientist reports — burny_tech · 2026-09-23
- Mathematician says OpenAI likely proved a significant partial case of the Hodge conjecture — burny_tech · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23