Kimi K3 Architecture Explained: Building a 2.8T Parameter Open Model via 'Active Forgetting'
AccBalanced · x · 2026-07-31
A detailed technical breakdown of the Kimi K3 model. K3 is currently the largest open-source model with 2.8 trillion total parameters, though only 104 billion are activated during inference.
- Core Innovation: To solve the slowdown caused by accumulating memory in long contexts, K3 introduces the Kimi Delta Attention mechanism.
- Active Forgetting: Three out of every four network layers use a fixed-size working memory, overwriting stale information to prevent unbounded memory growth. The fourth layer retains a compressed record for exact detail recovery.
- Validation: The team first validated this mechanism on a 48B parameter model, proving significant memory reduction before scaling up.
Related event: Inside Kimi K3: 2.8T Parameters and Active Forgetting(2 posts)→
More from Models
- Gemini V4 Rumored to Train on Tens of Trillions of Tokens for Multimodal — teortaxesTex · 2026-07-31
- DeepSeek V4-Flash Cracks Complex Russian Joke That Trips Up Other LLMs — teortaxesTex · 2026-07-31
- Model Selection is Becoming Org Design: Structuring AI Workflows — every · 2026-07-31
- DeepSeek-V4-Flash Benchmarks Leak: Higher IQ Density but Questionable Pricing — teortaxesTex · 2026-07-31
- Testing Inkling Small: A Vision-Equipped Model That Can Build Flappy Bird — LiTianleli · 2026-07-31
- DeepSeek-V4-Flash Official API Launches with Enhanced Agent Capabilities and Ultra-low Pricing — 智东西 · 2026-07-31