Kimi K3 Optimization: KDA Numerical Stability and Chunkwise Parallelism
nrehiew_ · x · 2026-07-29
Detailing the engineering optimizations for numerical stability in Kimi K3's KDA. In chunkwise parallel forms, large chunks cause decay and numerical instability. K3 mitigates this by further splitting computations into 16-token tiles and applying decay within each.
It also replaces softplus with a scaled sigmoid for the forget gate, lowering decay bounds to -5 and enabling Tensor Core usage for diagonal entries. Finally, the output gate is modified to be data-dependent and full-rank.
Related event: Kimi K3 report reveals training and systems stack(19 posts)→
More from Research
- Kimi K3 Architecture: KV Cache Offloading vs. KDA Recurrent State — zephyr_z9 · 2026-07-29
- Small-model orchestration roughly doubled task completion in a 100-task benchmark — _raydeStar · 2026-07-29
- Paper adds a human-only authorship attestation to a quantum matrix result — burny_tech · 2026-07-29
- Biohub is hiring for an AI wet-lab role to build biology models — proteinrosh · 2026-07-29
- An AI digest scans 92 journals every week and turns them into one RSS feed — Afinetheorem · 2026-07-29
- A weekly PDB-synced leaderboard tracks open cofolding models — rishabh16_ · 2026-07-29