vLLM to Natively Support Kimi K3 at Launch
AccBalanced · x · 2026-07-17
The vLLM team congratulated Moonshot AI on the release of Kimi K3 and announced a partnership between the two. Moonshot AI contributed the implementation code for KDA (Kimi Delta Attention) prefix caching to vLLM, breaking traditional prefix caching assumptions and enabling the open-source community to efficiently serve long-context inference on day one. vLLM will natively support Kimi K3 at launch, with the model weights scheduled to be open-sourced on July 27, 2026.
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21