Kimi K3 and Delta Attention
Gauri_the_great · x · 2026-07-17
The quoted content notes that Kimi K3 is a 2.8 trillion parameter model, making it one of the largest open-weight models released to date.
More notably, it introduces the concept of Attention Residuals / Delta Attention: instead of treating residuals merely as fixed identity pathways from the previous layer, each layer can selectively retrieve representations from multiple earlier layers. This turns the intermediate layers of deep networks into an addressable memory, where any layer can act as a "memory slot" for subsequent computations.
The poster expresses high anticipation for @rasbt's upcoming blog post explaining this architecture.
Related event: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(94 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11