vLLM Announces Day 0 Support for Kimi K3: Run 2.8T MoE on 8 B300 GPUs
vllm_project · x · 2026-07-30
The vLLM team announced efficient, production-scale Day 0 support for Moonshot AI's newly open-sourced Kimi K3 weights.
Architecture & Challenges
- Scale: Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with 16 active experts per token, a 1M-token context window, and native vision.
- Core Design: Built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes).
- Engineering: The team integrated KDA, MXFP4 MoE, KV cache management, prefill/decode disaggregation, speculative decoding, and long-context recipes into a runnable engine.
Deployment
The easiest deployment runs on 8 NVIDIA B300 or AMD MI355X GPUs. Users can also enable the open-sourced Inferact DSpark speculator via specific flags to boost inference speed.
Related event: vLLM and AMD Announce Day-0 Inference Support for Kimi K3(7 posts)→
More from Infra
- Microsoft Cloud Annual Revenue Tops $100B as AI Business Surges — tomwarren · 2026-07-30
- Vector DB Turbopuffer Adds Beta Support for Late Interaction Models — lateinteraction · 2026-07-30
- Qualcomm Q3 revenue beats estimates but weak Q4 EPS guide weighs — firstadopter · 2026-07-30
- Sam Altman Understands Why People Don't Want AI Data Centers in Their Backyards — businessinsider · 2026-07-30
- Zuckerberg: We're getting compute offers at a significant premium — firstadopter · 2026-07-30
- Dual GPU inference with RTX 4090 + 3060: speed impact and optimization tips — cosmoschtroumpf · 2026-07-30