Kimi K3’s constant-state design cuts long-context memory use by about 73%
bookwormengr · x · 2026-07-29
The post analyzes Kimi K3’s constant-state design and argues that its long-context memory savings may matter more than critics admit.
- The image compares memory use for 128K and 1M tokens: All-MLA (93L) uses 13.7 GB and 107 GB respectively, while Kimi K3 (24 MLA + KDA state) uses 3.5 GB + 0.22 GB and 29.0 GB + 0.22 GB.
- The core claim is that Kimi K3 avoids linearly growing KV cache for most of the conversation by using Kimi Delta Attention (KDA) / shared-key-value attention variants.
- The post frames this as a rebuttal to the “it’s already priced in” argument: if long-context workloads become much cheaper on memory, the pricing and demand dynamics for memory chips and related infrastructure could change materially.
- It also notes that the debate on X is not just about architecture novelty, but about whether long-context inference economics have been underestimated.
Related event: Kimi K3 Optimizes Long-Context: Massive VRAM Drop and 6x Speedup(2 posts)→
More from Infra
- Antirez says splitting large model files is an anti-pattern Hugging Face should hide — antirez · 2026-07-29
- Post claims GPU prices will rise 15% next month and local AI will be restricted — AIFlow_ML · 2026-07-29
- SK Hynix posts a mixed Q2 as revenue misses but net profit tops estimates — basedjensen · 2026-07-29
- vLLM maintainer wins AMD AI Core Contributor award at Advancing AI — vllm_project · 2026-07-29
- AI and chip stocks are starting to look like a parabolic run fading into distribution — Dan_Jeffries1 · 2026-07-29
- AI capital is still flowing, but smaller open-weights and inference bets may be easier to fund — vaibhavbetter · 2026-07-29