Kimi K3 Architecture: Dropping RoPE for KDA and LatentMoE
khademinori · x · 2026-07-30
The thread explores the key architectural trade-offs in the Kimi K3 model and how they translate to strong performance.
- Core Architecture: K3 drops RoPE entirely, even though its original author is on the team. Instead, it uses Kimi Delta Attention (KDA), which acts as a recurrent, content-addressable matrix-state RNN, making tokens order-aware before global MLA.
- New Components: The main addition compared to Kimi Linear is the LatentMoE, which compresses large linear layers via down-projection, similar to Nemotron 3 Ultra.
- Scaling Success: The architecture of Kimi Linear was successfully scaled from 48B to a 2.8T production-scale frontier model, proving its stability under extreme scaling.
Related event: Inside Kimi K3: 2.8T Parameters and Active Forgetting(2 posts)→
More from Models
- KOL Rebuts 'DeepSeek Missed the Agent Boat' Claim, Urges Real-World Testing — teortaxesTex · 2026-07-31
- DeepSeek Praised as the Only Force Driving Down LLM API Prices — teortaxesTex · 2026-07-31
- Redis Creator antirez Advances Locally Runnable DwarfStar Model — antirez · 2026-07-31
- SOTA LLMs Tend to Reword Text When Asked Only to Fix Typos — mitsuhiko · 2026-07-31
- Alibaba Releases Qwen-Audio-3.0-ASR-Flash with Enhanced Hotwords and Domain Recognition — Alibaba_Qwen · 2026-07-31
- Luna Model Prices Slashed by 80%, Ushering in Cheap Intelligence Era — teortaxesTex · 2026-07-31