Kimi K2 Architecture: Mixing KDA and Gated MLA, and the Forget-Gate Mechanism
suchenzang · x · 2026-07-28
The author analyzes Kimi K2's attention scaling, noting it mixes local KDA (Kimi Delta Attn) with global Gated MLA (Deepseek V2) in a 3:1 ratio. KDA is constrained with qk-norm, and channel-wise decay is framed as a 'forget-gate,' acting similarly to weight decay to constrain dynamic range.
More from Research
- Kimi K3 bounds decay at -5 to keep chunkwise KDA inside BF16 range — suchenzang · 2026-07-28
- Block Attention Residuals cuts attention overhead from O(Ld) to O(Nd) — stochasticchasm · 2026-07-28
- Celltype pitches LLMs that predict biological response and is hiring in New York — david_van_dijk · 2026-07-28
- Claude Opus 5 keeps Opus 4.8 pricing while matching Fable 5 within 0.5% on coding — AlexKim · 2026-07-28
- MLA Architecture Details: Will Full-Rank Gate Projection Cause Parameter Explosion? — stochasticchasm · 2026-07-28
- Kimi K3 replaces KDA’s low-rank output gate with an input-dependent full-rank projection — stochasticchasm · 2026-07-28