vLLM ships day-0 serving support for Moonshot’s 2.8T-parameter Kimi K3
vllm_project · x · 2026-07-28
vLLM publishes a day-0 serving guide for Moonshot’s Kimi K3
vLLM says it now supports Kimi K3 on day 0 as Moonshot’s weights are public. The post is a practical deployment guide for running the model in production.
Key details from the guide:
- Kimi K3 is a 2.8T-parameter MoE model with 16 of 896 experts active per token.
- It uses Kimi Delta Attention (KDA) and Attention Residuals (AttnRes).
- The model has a 1M-token context window and native vision.
- The engineering challenge is making KDA, MXFP4 MoE execution, KV cache management, prefill/decode disaggregation, speculative decoding, and long-context serving work together.
The attached diagrams focus on cache layout, decode throughput, and the memory/performance tradeoffs of the model’s architecture.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23