vLLM Announces Day-0 Support for Kimi K3: Run 2.8T MoE on 8 B300 GPUs
vllm_project · x · 2026-07-30
vLLM officially announced efficient Day-0 production support for Moonshot AI's open-weight Kimi K3 model. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts (MoE) model with 16 active experts per token, featuring a 1M-token context window and native vision capabilities.
The post provides a detailed deployment guide, noting that the easiest way to run the model is using 8 NVIDIA B300 GPUs or 8 AMD MI355X GPUs. Additionally, Inferact has trained and open-sourced a DSpark speculative decoder specifically for Kimi K3, which can be enabled via specific serve arguments to boost inference speed. Baseten has also launched a fast API service powered by vLLM, allowing users to leverage the model without managing infrastructure.
Related event: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(14 posts)→
More from Infra
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23