DigitalOcean and vLLM Detail Day-0 Inference Recipe for 2.8T Param Kimi K3
vllm_project · x · 2026-07-31
DigitalOcean, in collaboration with the vLLM team, successfully deployed the massive 2.8 trillion parameter Kimi K3 model into production on day zero.
Hardware & Architecture
- Selected NVIDIA HGX B300 and AMD Instinct MI350X GPUs, providing 288GB VRAM per 8-GPU node.
- K3 weights consume 1.56 TB (approx. 195 GiB per GPU), leaving a practical headroom for KV cache and activations.
- Utilized llm-d for the distributed inference stack to natively support GPU heterogeneity across AMD and NVIDIA platforms.
Model Complexity
Features 896 routed experts and an interleaved attention stack combining 69 Kimi Delta Attention (KDA) layers and 24 MLA layers.
Related event: vLLM Day-0 Support for Kimi K3: Run 2.8T Model on 8 B300 GPUs(14 posts)→
More from Infra
- Yahoo Taiwan Ecommerce Cuts Listing Times to 2 Minutes with On-Device AI — gaganghotra_ · 2026-07-31
- xAI to Remove All Temporary Turbines by 2027, Transitioning to 1.2GW Permanent Power — XFreeze · 2026-07-31
- Run Local AI Agents on Consumer GPUs: Cloud Planning + Local Execution — fire_inabottle · 2026-07-31
- Microsoft Hits Record Single-Day Market Cap Gain as AI Demand Accelerates — firstadopter · 2026-07-31
- Ghibli Trend Triggers Massive AI Compute Shortage for Inference — BenBajarin · 2026-07-31
- Qualcomm Teases Two Flagship Mobile Chips at Snapdragon Summit — ryanshrout · 2026-07-31