Kimi K3 and Delta Attention

Gauri_the_great · x · 2026-07-17

The quoted content notes that Kimi K3 is a 2.8 trillion parameter model, making it one of the largest open-weight models released to date.

More notably, it introduces the concept of Attention Residuals / Delta Attention: instead of treating residuals merely as fixed identity pathways from the previous layer, each layer can selectively retrieve representations from multiple earlier layers. This turns the intermediate layers of deep networks into an addressable memory, where any layer can act as a "memory slot" for subsequent computations.

The poster expresses high anticipation for @rasbt's upcoming blog post explaining this architecture.

Related event: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(94 posts)→

Original post →

More from Models

Models channel →