DigitalOcean and vLLM Detail Day-0 Inference Recipe for 2.8T Param Kimi K3

vllm_project · x · 2026-07-31

DigitalOcean, in collaboration with the vLLM team, successfully deployed the massive 2.8 trillion parameter Kimi K3 model into production on day zero.

Hardware & Architecture

Model Complexity

Features 896 routed experts and an interleaved attention stack combining 69 Kimi Delta Attention (KDA) layers and 24 MLA layers.

Related event: vLLM Day-0 Support for Kimi K3: Run 2.8T Model on 8 B300 GPUs(14 posts)→

Original post →

More from Infra

Infra channel →