Kimi K2 Architecture: Mixing KDA and Gated MLA, and the Forget-Gate Mechanism

suchenzang · x · 2026-07-28

The author analyzes Kimi K2's attention scaling, noting it mixes local KDA (Kimi Delta Attn) with global Gated MLA (Deepseek V2) in a 3:1 ratio. KDA is constrained with qk-norm, and channel-wise decay is framed as a 'forget-gate,' acting similarly to weight decay to constrain dynamic range.

Original post →

More from Research

Research channel →