DigitalOcean and vLLM Detail Day-0 Inference Recipe for 2.8T Param Kimi K3
vllm_project · x · 2026-07-31
DigitalOcean, in collaboration with the vLLM team, successfully deployed the massive 2.8 trillion parameter Kimi K3 model into production on day zero.
Hardware & Architecture
- Selected NVIDIA HGX B300 and AMD Instinct MI350X GPUs, providing 288GB VRAM per 8-GPU node.
- K3 weights consume 1.56 TB (approx. 195 GiB per GPU), leaving a practical headroom for KV cache and activations.
- Utilized llm-d for the distributed inference stack to natively support GPU heterogeneity across AMD and NVIDIA platforms.
Model Complexity
Features 896 routed experts and an interleaved attention stack combining 69 Kimi Delta Attention (KDA) layers and 24 MLA layers.
Related event: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(14 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23