vLLM Releases Kimi K3 Deployment Guide: Supports B300 and MI355X
vllm_project · x · 2026-07-30
The vLLM team released a detailed technical guide for Day-0 support of the Kimi K3 model. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with 16 active experts per token, featuring a 1M-token context window and native vision.
The post dives into how vLLM adapts to Kimi K3's unique KDA and AttnRes architecture, tackling deployment challenges like MXFP4 MoE, KV cache management, and prefill/decode disaggregation. It recommends running on 8 NVIDIA B300 or AMD MI355X GPUs and supports open-source DSpark speculative decoding acceleration provided by Inferact.
Related event: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(14 posts)→
More from Infra
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23