Kimi K3 and Delta Attention
Gauri_the_great · x · 2026-07-17
The quoted content notes that Kimi K3 is a 2.8 trillion parameter model, making it one of the largest open-weight models released to date.
More notably, it introduces the concept of Attention Residuals / Delta Attention: instead of treating residuals merely as fixed identity pathways from the previous layer, each layer can selectively retrieve representations from multiple earlier layers. This turns the intermediate layers of deep networks into an addressable memory, where any layer can act as a "memory slot" for subsequent computations.
The poster expresses high anticipation for @rasbt's upcoming blog post explaining this architecture.
Related event: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(94 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22