Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds
Moonshot AI open-sourced Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, with vLLM providing production-grade inference support on day one, requiring as few as 8 B300 GPUs to run the full model and natively supporting AMD ROCm. The model is now available on DigitalOcean, Modal, and Baseten, enabling rapid deployment of production APIs. This significantly lowers the barrier to deploying massive open-source models and breaks the single-hardware-ecosystem limitation.
Confirmed
- Model architecture: Kimi K3 has 2.8T total parameters, MoE architecture, activates 16 experts per token, supports 1M token context window and native vision capabilities.
- Hardware deployment: vLLM confirmed the model can run with as few as 8 B300 GPUs, and supports NVIDIA Grace Blackwell, Blackwell, Hopper, and NVL72 systems, integrating Dynamo for efficient operation.
- AMD ecosystem support: vLLM natively supports AMD ROCm, allowing full model deployment on Instinct series hardware. AMD executive Anush Elangovan confirmed day-0 support for Kimi K3 and MI355X hardware through collaboration, emphasizing "speed is the moat."
- Cloud deployment: Kimi K3 is available on DigitalOcean, Modal, and Baseten, with vLLM providing the underlying inference service, enabling developers to quickly deploy production-grade API endpoints.
- Inference performance: TensorWave noted that through AMD ecosystem adaptation, RadixArk's DSPARK achieved 423 TPS throughput running the model.
Why it matters
- Lower deployment barrier: The 2.8T parameter model has high VRAM requirements; vLLM's efficient support allows running with just 8 B300s, greatly reducing compute costs for massive models.
- Breaking hardware monopoly: Day-0 native support for AMD ROCm provides developers with a strong alternative to NVIDIA, accelerating diversification of AI inference hardware ecosystem.
2026-07-30 ~ 2026-07-31 · 14 related posts
- Episode 1: vLLM brings day-0 support to Moonshot’s Kimi K3(2026-07-27, 11 posts)
- Episode 2: Kimi K3 Now Available for Inference and Fine-Tuning on Fireworks(2026-07-28, 2 posts)
- Episode 3: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(2026-07-28, 8 posts)
- Episode 4: Kimi K3 Open-Weight Release Sparks Debate on Open Source and Infrastructure(2026-07-28, 5 posts)
- Episode 5: Kimi K3 Self-Hosting Can Break Even in Under 100 Days(2026-07-28, 4 posts)
- Episode 6: Tinkering with Local Quantized K3 Inference on Mac Hardware(2026-07-29, 2 posts)
- Episode 7: Kimi K3 Open Weights Demand Data Center Hardware(2026-07-29, 3 posts)
- Episode 8: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2026-07-29, 2 posts)
- Episode 9: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(2026-07-30, 14 posts)
Primary sources
- Kimi K3 runs on vLLM + AMD from Day 0, supporting 2.8T params on Instinct — vllm_project ·
- vLLM Announces Day 0 Support for Kimi K3: Run 2.8T MoE on 8 B300 GPUs — vllm_project ·
- AMD MI355X Achieves Day 0 Support for Kimi K3, Boosting AI Ecosystem — AnushElangovan ·
- [source] Kimi K3 runs on vLLM + AMD from Day 0, supporting 2.8T params on Instinct — vllm_project · 2026-07-30
- vLLM Announces Day-0 Support for Kimi K3: Deploying the 2.8T Parameter Model — vllm_project · 2026-07-30
- Kimi K3 Lands on DigitalOcean Powered by vLLM for Efficient Inference — vllm_project · 2026-07-30
- [source] AMD MI355X Achieves Day 0 Support for Kimi K3, Boosting AI Ecosystem — AnushElangovan · 2026-07-30
- vLLM Announces Day 0 Support for Kimi K3 Across NVIDIA Architectures — vllm_project · 2026-07-30
- vLLM and Modal Announce Day-0 Support for Kimi K3 Deployment — vllm_project · 2026-07-30
- vLLM Releases Kimi K3 Deployment Guide: Supports B300 and MI355X — vllm_project · 2026-07-30
- Kimi K3 Available on Baseten with vLLM-Powered Production API — vllm_project · 2026-07-30
- Kimi K3 Launches with Day 0 Support on AMD Instinct via vLLM — vllm_project · 2026-07-30
- DigitalOcean and vLLM Detail Day-0 Inference Recipe for 2.8T Param Kimi K3 — vllm_project · 2026-07-31
- Kimi K3 open-source model lands on AMD ecosystem, RadixArk hits 423 TPS — BanghuaZ · 2026-07-31
3 near-duplicate retellings: vllm_project · vllm_project · vllm_project