vLLM adds day-zero support for Kimi K3, a 2.8T MoE model with 1M context
AAAzzam · x · 2026-07-27
vLLM adds day-zero support for Kimi K3
The vLLM project says Kimi K3 can be served on its stack as soon as the weights are public.
- Kimi K3 is described as a 2.8T-parameter Mixture-of-Experts model.
- It uses 16 of 896 experts per token.
- The model has a 1M-token context window and native vision understanding.
- vLLM highlights Kimi Delta Attention, a hybrid attention design intended to make million-token context affordable.
- The team also thanked Moonshot, inferact, Nvidia, AMD, and the vLLM community for the launch integration.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23