vLLM adds day-0 support for Moonshot’s 2.8T-parameter Kimi K3
woosuk_k · x · 2026-07-28
vLLM announced day-0 support for Moonshot AI’s Kimi K3, saying the model can be served as soon as the weights are public.
The post highlights K3’s scale and the supporting engineering work around prefix caching, DSpark training, and kernel optimizations. K3 is described as a 2.8T-parameter MoE model with a 1M-token context window, native multimodal understanding, and Kimi Delta Attention to make long-context inference more affordable.
Related event: vLLM Announces Day-0 Inference Support for 2.8T Parameter Kimi K3(10 posts)→
More from Infra
- Calibrated Qwen3.6-27B quantization tests weight groups before compressing them — enginetown · 2026-07-28
- Frozen 12B system reuses verified memory at zero tokens and 6,000,000-token context — Corbenci · 2026-07-28
- OpenAI’s expected $750 billion compute spend puts Anthropic’s leasing strategy under pressure — remybigot · 2026-07-28
- AMD, SGLang and Moonshot ship together as the chip wars shift to infrastructure — AnushElangovan · 2026-07-28
- A user says GPT-5.4, Opus 4.6, and Kimi k3 already cover most needs — haider1 · 2026-07-28
- Fast-plate-ocr adds lightweight license plate recognition with Keras 3 and ONNX — tom_doerr · 2026-07-28