vLLM Announces Day 0 Support for Kimi K3 Across NVIDIA Architectures
vllm_project · x · 2026-07-30
The vLLM project has announced first-class support for the Kimi K3 model, optimized for NVIDIA hardware.
The deployment serves K3 across Grace Blackwell, Blackwell, Hopper, and NVL72 systems. It also integrates with Dynamo to maintain speed and efficiency for such a large model at production scale. This support was available on Day 0, the moment the model weights went public.
Related event: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(14 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23