Kimi K3 replaces KDA’s low-rank output gate with an input-dependent full-rank projection
stochasticchasm · x · 2026-07-28
- Another quote from the Kimi K3 technical report discusses the model’s full-rank gate.
- The report says Kimi K3 changes KDA’s output gate from the low-rank setup used by Kimi Linear to an input-dependent full-rank projection.
- After applying head-wise RMSNorm, KDA uses data-dependent output gating, with the gate written as a sigmoid of a matrix projection applied to the token input.
More from Research
- Kimi K3 report introduces SiTU-GLU, a bounded tanh × sigmoid gate that approximates SwiGLU — stochasticchasm · 2026-07-28
- Kimi K3’s MoE routing may be driving higher expert-parallel communication costs — stochasticchasm · 2026-07-28
- Kimi K3 finds 16 new vulnerabilities and beats GLM-5.2 on an exploit benchmark — zephyr_z9 · 2026-07-28
- Kimi K3 bounds decay at -5 to keep chunkwise KDA inside BF16 range — suchenzang · 2026-07-28
- Block Attention Residuals cuts attention overhead from O(Ld) to O(Nd) — stochasticchasm · 2026-07-28
- MLA Architecture Details: Will Full-Rank Gate Projection Cause Parameter Explosion? — stochasticchasm · 2026-07-28