A reading order for Kimi K3 traces the papers behind Moonshot’s frontier model
East-Muffin-6472 · reddit · 2026-07-29
A Reddit post lays out a reading order for understanding Kimi K3 from first principles.
- It starts with Linear Transformers Are Secretly Fast Weight Programmers as the conceptual foundation for modern linear attention.
- It then moves to Gated DeltaNet, Kimi Linear / Kimi Delta Attention (KDA), LatentMoE → Stable LatentMoE, and Attention Residuals.
- The post explains that Kimi K3 combines several research threads: efficient linear attention, better state updates, sparse expert routing, and improved residual mixing.
- It also recommends reading the Kimi model evolution in order: K1.5 → K2 → K2.5 → K3 to see how the architecture and training stack evolved over time.
Related event: Community Outlines Research Roadmap Behind Kimi K3(2 posts)→
More from Research
- Four months of Gabor-wavelet image generation led to a broader Claude-assisted research project — pixlpa · 2026-07-30
- Agent Retrieval Bench shows coding agents often fail before patch generation even starts — Bowen Qin · 2026-07-30
- When MCP servers drift, user settings may stop applying to the capability you meant — Loocor · 2026-07-30
- AI-generated Lean proof of Collatz solution exploited a kernel bug — rbhar90 · 2026-07-30
- A new robot imitation method learns 1,000 tasks in under 24 hours of demo time — chris_j_paxton · 2026-07-30
- Kimi K3 Architecture: KV Cache Offloading vs. KDA Recurrent State — zephyr_z9 · 2026-07-29