Kimi K3 claims 2.5× better scaling efficiency with a three-axis architecture
suchenzang · x · 2026-07-28
The Kimi K3 architecture scales information flow along three axes: sequence length, depth, and width.
- Sequence length: Hybrid Attention combines three Kimi Delta Attention layers with one Gated MLA layer per block to support long-context token mixing.
- Depth: Attention Residuals let modules retrieve representations from the embedding, current block, and previous blocks.
- Width: A Stable LatentMoE layer performs sparse channel mixing, activating 16 of 896 routed experts per token.
- Multimodal path: MoonViT-V2 encodes images and videos, then a lightweight projector maps features into the shared embedding space.
- Result: With refined training/data recipes and Per-Head Muon, the paper claims about 2.5× better scaling efficiency than Kimi K2.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Research
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23