Deep Dive into Kimi K3 Architecture: 2.8T Parameters and Attention Innovations
Kimi K3(2.8T 参数、104B 激活参数的开权重 MoE 模型)发布后,技术社区对其架构进行了深度复盘,确认其并非单纯的参数堆砌,而是注意力机制与底层架构的重大演进。Sebastian Raschka 指出,K3 是 Kimi Linear 的大规模生产化版本,层数从 27 层扩至 93 层,并用潜 MoE 替代了标准 MoE。
已确认
- 核心架构参数:模型总参数 2.8T,激活参数 104B,共 93 层。其后训练配方分为 SFT cold start 等三个主要步骤。
- 注意力机制革新:K3 彻底去除了位置编码/位置嵌入,偏离了传统 Transformer 规范。它采用了混合的线性注意力与全注意力设计,并使用了自研的 Kimi Linear 变体。
- 长上下文优化:核心组件 Kimi Delta Attention (KDA) 极大优化了长上下文推理的显存成本。数据显示,在 128K tokens 下,其常量状态设计相比 All-MLA 显存降低约 73%;KDA 整体将 KV cache 压低 75%,并实现 6 倍提速。
为什么重要
多位分析者(如 peterjliu 和 bookwormengr)强调,理解 Kimi K3 不能仅看规模,而应将其置于从 GPT-2 以来的长期缩放与注意力原语演进史中。它证明了通过架构层面的创新(如去除位置编码、引入 KDA),可以在不盲目扩张算力的前提下,有效突破长上下文等推理瓶颈。
2026-07-27 ~ 2026-07-29 · 16 related posts
- Episode 1: Rumor: Gemini 3.5 Performance Rivals GPT-5.5(2026-07-05, 3 posts)
- Episode 2: Rumored Release Schedule for Frontier AI Models in July(2026-07-06, 5 posts)
- Episode 3: Multiple Major AI Models Set for Dense Release(2026-07-08, 3 posts)
- Episode 4: Gemini 3.5 Pro Faces Multiple Delay Rumors and Performance Scrutiny(2026-07-10, 6 posts)
- Episode 5: AI Infrastructure Boom: Open Source vs Frontier Models(2026-07-13, 3 posts)
- Episode 6: AI Efficiency Gains May Amplify Demand(2026-07-13, 2 posts)
- Episode 7: AI Safety Focus Shifts from Model Output to Agent Execution Risks(2026-07-13, 9 posts)
- Episode 8: Rumored Gemini 3.5 Pro Launch Nears(2026-07-14, 3 posts)
- Episode 9: Kimi K3 hype builds as KIVINE appears on Arena(2026-07-14, 43 posts)
- Episode 10: Rumors Grow of Another Gemini 3.5 Pro Delay(2026-07-15, 7 posts)
- Episode 11: Wave of Frontier AI Model Releases Imminent(2026-07-15, 2 posts)
- Episode 12: The Open Source AI Debate: Security, Research, and Monopoly(2026-07-15, 10 posts)
- Episode 13: Wave of new model release rumors surfaces, none yet confirmed(2026-07-15, 7 posts)
- Episode 14: Kimi K3 Debuts Strong, Narrowing the Open-Weight Gap(2026-07-15, 184 posts)
- Episode 15: Kimi K3 Tops Frontend Code Arena and Sparks Debate(2026-07-16, 53 posts)
- Episode 16: AI Frontier Advantage Narrows to Months(2026-07-16, 2 posts)
- Episode 17: Kimi K3 Triggers a Reassessment of Chinese Frontier AI(2026-07-16, 94 posts)
- Episode 18: Kimi K3 Sparks AI Community Buzz with Top-Tier Performance(2026-07-16, 3 posts)
- Episode 19: Kimi K3 Sparks Debate Over Real-World Coding Ability(2026-07-16, 6 posts)
- Episode 20: Kimi K3 Sparks Debate Over Open-Weight Frontier AI(2026-07-17, 15 posts)
Primary sources
- [source] Kimi Delta Attention cuts KV cache by 75% and speeds million-token decoding by 6× — johnseach · 2026-07-27
- A 7-year tour of open-model architecture explains why Kimi K3 is not just bigger — philipkiely · 2026-07-27
- A deep dive traces Kimi K3’s lineage back to GPT-2 across 8 papers and 48 hours — Madisonkanna · 2026-07-27
- A long history of LLM architectures leads from GPT-2 to Kimi K3 — baseten · 2026-07-28
- Kimi K3 paper drops position embeddings and pushes beyond Transformer orthodoxy — peterjliu · 2026-07-28
- Kimi K3 reportedly keeps block attention residuals in a 93-layer design — burny_tech · 2026-07-28
- Kimi K3 is framed as far from a Transformer in a new attention-primitive overview — AccBalanced · 2026-07-28
- Kimi K3’s architecture is explained as an industrial plumbing system — doodlestein · 2026-07-28
- Baseten engineer says Kimi K3’s leap came from a chain of targeted fixes, not scale alone — khademinori · 2026-07-28
- [source] Kimi K3 scales Kimi Linear to 2.8T parameters and drops RoPE for NoPE — rasbt · 2026-07-28
- Kimi K3 paper points to hybrid attention, Kimi Linear, and no positional embeddings — peterjliu · 2026-07-28
- MoonshotAI’s FlashKDA commit hints at a hybrid attention mainline model — peterjliu · 2026-07-29
- Kimi K3 report says a 2.8T MoE used RL experts and multi-teacher distillation — cwolferesearch · 2026-07-29
- Kimi K3 Weights Released: 93 Layers, Latent MoE Replaces Standard MoE — hugobowne · 2026-07-29
- [source] Kimi K3’s constant-state design cuts long-context memory use by about 73% — bookwormengr · 2026-07-29
1 near-duplicate retellings: bookwormengr