vLLM to Natively Support Kimi K3 at Launch

AccBalanced · x · 2026-07-17

The vLLM team congratulated Moonshot AI on the release of Kimi K3 and announced a partnership between the two. Moonshot AI contributed the implementation code for KDA (Kimi Delta Attention) prefix caching to vLLM, breaking traditional prefix caching assumptions and enabling the open-source community to efficiently serve long-context inference on day one. vLLM will natively support Kimi K3 at launch, with the model weights scheduled to be open-sourced on July 27, 2026.

Original post →

More from Infra

Infra channel →