vLLM Announces Day-0 Support for Kimi K3: Run 2.8T MoE on 8 B300 GPUs
vllm_project · x · 2026-07-30
vLLM officially announced efficient Day-0 production support for Moonshot AI's open-weight Kimi K3 model. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts (MoE) model with 16 active experts per token, featuring a 1M-token context window and native vision capabilities.
The post provides a detailed deployment guide, noting that the easiest way to run the model is using 8 NVIDIA B300 GPUs or 8 AMD MI355X GPUs. Additionally, Inferact has trained and open-sourced a DSpark speculative decoder specifically for Kimi K3, which can be enabled via specific serve arguments to boost inference speed. Baseten has also launched a fast API service powered by vLLM, allowing users to leverage the model without managing infrastructure.
Related event: Kimi K3 Open-Sourced with Day-0 Native Support from vLLM and AMD(11 posts)→
More from Infra
- QuixiAI Open-Sources SlimServe: Fast Inference for GLM on AMD MI300X — QuixiAI · 2026-07-30
- ThunderAgent by Together AI: 2x Faster Agentic Inference (ICML 2026) — togethercompute · 2026-07-30
- Together Compute unveils multi-node inference engine with near-linear scaling, adopted by SkyRL and NVIDIA Dynamo — togethercompute · 2026-07-30
- DigitalOcean Becomes Day 0 Inference Partner for Kimi K3 with 1M Context — vllm_project · 2026-07-30
- Kimi K3 Launches with Day 0 Support on AMD Instinct via vLLM — vllm_project · 2026-07-30
- Debunking the DeepSeek and Chinese Lithography Panic: Exaggerated Costs and Gaps — teortaxesTex · 2026-07-30