vLLM Day-0 Support for Kimi K3: Deploying a 2.8T Parameter MoE Model
vllm_project · x · 2026-07-30
The vLLM team announced efficient Day-0 support for Moonshot AI's open-weight Kimi K3 model. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts (MoE) model with 16 experts active per token, featuring a 1M-token context window and native vision capabilities.
The post details the engineering challenges and solutions for adapting vLLM to this architecture:
- Architecture Adaptation: Specialized kernel and cache optimizations were made for new architectural features like Kimi Delta Attention (KDA), Attention Residuals (AttnRes), and MXFP4 MoE.
- Inference Acceleration: Inferact has open-sourced a DSpark speculator, which, combined with prefill/decode disaggregation and long-context deployment recipes, significantly boosts inference efficiency.
- Deployment Guide: Provides complete Docker startup commands and configurations for running the model on 8 NVIDIA B300 or AMD MI355X GPUs.
Related event: vLLM and AMD Announce Day-0 Inference Support for Kimi K3(7 posts)→
More from Infra
- US Commerce Dept Allocates $874M to Accelerate Semiconductor R&D — imjustnewatai · 2026-07-30
- Qualcomm Q3 revenue beats estimates but weak Q4 EPS guide weighs — firstadopter · 2026-07-30
- Sam Altman Understands Why People Don't Want AI Data Centers in Their Backyards — businessinsider · 2026-07-30
- Zuckerberg: We're getting compute offers at a significant premium — firstadopter · 2026-07-30
- Dual GPU inference with RTX 4090 + 3060: speed impact and optimization tips — cosmoschtroumpf · 2026-07-30
- Future 100T Param Model to Cost >$250B to Train, Says Joseph Jacks — JosephJacks_ · 2026-07-30