vLLM Releases Kimi K3 Deployment Guide: Supports B300 and MI355X
vllm_project · x · 2026-07-30
The vLLM team released a detailed technical guide for Day-0 support of the Kimi K3 model. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with 16 active experts per token, featuring a 1M-token context window and native vision.
The post dives into how vLLM adapts to Kimi K3's unique KDA and AttnRes architecture, tackling deployment challenges like MXFP4 MoE, KV cache management, and prefill/decode disaggregation. It recommends running on 8 NVIDIA B300 or AMD MI355X GPUs and supports open-source DSpark speculative decoding acceleration provided by Inferact.
Related event: Kimi K3 Open-Sourced with Day-0 Native Support from vLLM and AMD(11 posts)→
More from Infra
- DigitalOcean Becomes Day 0 Inference Partner for Kimi K3 with 1M Context — vllm_project · 2026-07-30
- Kimi K3 Launches with Day 0 Support on AMD Instinct via vLLM — vllm_project · 2026-07-30
- Debunking the DeepSeek and Chinese Lithography Panic: Exaggerated Costs and Gaps — teortaxesTex · 2026-07-30
- Running Kimi K3 on CPU: Custom Q3 Quantization Takes 1.1TB, Hits 4.2 t/s — Fun-Meaning-6474 · 2026-07-30
- GPT-6 Expected to Autonomously Optimize Its Own Inference Compute — imjustnewatai · 2026-07-30
- Kimi K3 Available on Baseten with vLLM-Powered Production API — vllm_project · 2026-07-30