KDA’s low-rank gate doesn’t fix attention sink, author says after scaling it up
stochasticchasm · x · 2026-07-28
- The author says KDA scaling moved the output gate from a low-rank design to a full-rank projection.
- They note that the original justification was to address the attention-sink issue, but that rationale appears to have been removed.
- Their conclusion is blunt: low-rank does not solve attention sink, and it adds more problems at scale.
- The attached figure shows the paper text describing the gate change and related normalization/gating details.
Related event: Kimi K3 Architecture: Introducing SiTU-GLU and Full-Rank Projection(3 posts)→
More from Research
- Open-source ModPack turns robot teleoperation into a modular stack — SongShuran · 2026-07-28
- Kimi K3’s QB routing uses quantiles to balance MoE expert load — stochasticchasm · 2026-07-28
- Kimi K3 report adds NoPE, full-rank gating and FP32 attention training — suchenzang · 2026-07-28
- ELEMENTA is a large-scale DFT dataset for UMLIP and foundation-model training — xie_tian · 2026-07-28
- Coding agents are already reshaping empirical social science research — soumitrashukla9 · 2026-07-28
- Kimi K3 paper details SiTU-GLU and quantile balancing for 896-expert MoE — KyeGomezB · 2026-07-28