vLLM says K3’s 2.8T parameters were easy; caching them was the hard part
vllm_project · x · 2026-07-27
vLLM explains that fitting K3’s 2.8T parameters was not the hard part; caching them was.
Because most K3 layers use fixed-size KDA state instead of a growing KV cache, there is no per-token KV to hash. Prefix caching had to be redesigned around that architecture, and the post says every hybrid model after K3 inherits the result.
The attached diagram also shows how K3 combines components such as Stable LatentMoE, Gated MLA, and KDA inside a multi-block stack.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23