vLLM Announces Day-0 Support for Kimi K3: Deploying the 2.8T Parameter Model

vllm_project · x · 2026-07-30

vLLM officially announced efficient Day-0 support for Moonshot AI's newly released Kimi K3. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model featuring a 1M-token context window and native vision capabilities.

The blog post details the engineering efforts to adapt vLLM to Kimi K3's architecture, making KDA, MXFP4 MoE, KV cache management, prefill/decode disaggregation, and speculative decoding work together. The team provided a quick-start command, recommending 8 NVIDIA B300 or AMD MI355X GPUs for deployment, along with support for an open-sourced DSpark speculator by Inferact.

Related event: vLLM and AMD Announce Day-0 Inference Support for Kimi K3(7 posts)→

Original post →

More from Infra

Infra channel →