vLLM lists Kimi K3 serving options with disaggregation, KV offload and MoE backends
vllm_project · x · 2026-07-27
- The post lists the main serving features that can be enabled for Kimi K3 on vLLM.
- Highlights include prefill/decode disaggregation, topology-aware MoE backends (megamoe for EP, trtllm for TP>1), KV offloading, and support for tool calling, reasoning, and structured output.
- The attached screenshot shows the vLLM Recipes page for moonshotai/Kimi-K3, including multi-node tensor parallel deployment, hardware targets, and latency-oriented serving settings.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23