Kimi K3 switches KDA to a full-rank gate, as low-rank is said to miss the attention-sink fix
suchenzang · x · 2026-07-28
- The author says scaling Kimi K3 changed KDA’s output gate from the low-rank parameterization used in Kimi Linear to a full-rank, input-dependent projection.
- They note that the earlier low-rank version had been justified as a way to help with the attention-sink problem.
- In the new writeup, that justification is said to be gone, and the author argues the low-rank design did not actually solve attention sink.
- The attached paper excerpt also highlights the move toward lower-bounded decay and numerical-stability fixes using log-space reasoning and sigmoid truncation.
Related event: Kimi K3 Optimizes KDA: Prevents Overflow and Upgrades Gate(2 posts)→
More from Research
- Nature study says AI-redesigned protein starting points improve enzyme evolution — i_dg23 · 2026-07-28
- Open-source ModPack turns robot teleoperation into a modular stack — SongShuran · 2026-07-28
- ELEMENTA is a large-scale DFT dataset for UMLIP and foundation-model training — xie_tian · 2026-07-28
- Coding agents are already reshaping empirical social science research — soumitrashukla9 · 2026-07-28
- Kimi K3 paper details SiTU-GLU and quantile balancing for 896-expert MoE — KyeGomezB · 2026-07-28
- Aaron Hertzmann says the brain is not a computer, raising the stakes for AGI — yacineMTB · 2026-07-28