vLLM Officially Supports Kimi K3 Deployment, Requires 8x GB300 Minimum
vllm_project · x · 2026-08-07
vLLM announces it is officially verified by Moonshot to deliver performant inference for the Kimi K3 model.
Model & Deployment Details:
- Architecture: Kimi K3 is a 2.8T-parameter native multimodal MoE model with 16 active experts per token (out of 896). It utilizes Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), supporting a 1M-token context window and native vision.
- Hardware: Requires at least 8x GB300 (multi-node for production) or 8x MI355X/MI350X (ROCm).
- Environment: Ships with a dedicated Docker image requiring CUDA 13 (cu130) and r580+ NVIDIA drivers.
- Optimizations: Supports TP8, TEP16, and disaggregated P/D profiles. Recommends deepepv2 backend for cross-node communication.
More from Infra
- High GPU Temps Running Local AI Models Prompts Cooling Advice — Inner_West_4997 · 2026-08-07
- How a McDonald's Potato Supplier Became the Savior of America's DRAM Industry — DynamicWebPaige · 2026-08-07
- NVIDIA Open-Sources Full Local Speech Stack: ASR, TTS, and Codecs — ImaginaryRea1ity · 2026-08-07
- Hyperscalers Pivot to Behind-the-Meter Power to Bypass Grid Bottlenecks — BenBajarin · 2026-08-07
- Big Tech's 2026 AI Capex Hits $732.5B, Putting $1 Trillion in 2027 Within Reach — Beth_Kindig · 2026-08-07
- Analysis: AI Optical Implementations Remain Lumpy and Bespoke per Customer — BenBajarin · 2026-08-07