Moonshot’s Kimi K3 long-context design draws skepticism over 1M-token scaling
teortaxesTex · x · 2026-07-21
A thread questions whether Moonshot’s long-context research path is the right one, arguing that its post-R1 shift to MLA may be weak beyond 262K tokens and asking whether KDA can scale to 1M as well as CSA+HCA.
The attached architecture image shows Kimi K3’s design stack, including Stable LatentMoE, Gated MLA, KDA, and the Kimi Delta Attention backbone. The discussion frames long context as the AI era’s version of RAM and suggests the real moat may be organizational as much as model-level.
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11