vLLM adds day-zero support for Kimi K3, a 2.8T MoE model with 1M context
AAAzzam · x · 2026-07-27
vLLM adds day-zero support for Kimi K3
The vLLM project says Kimi K3 can be served on its stack as soon as the weights are public.
- Kimi K3 is described as a 2.8T-parameter Mixture-of-Experts model.
- It uses 16 of 896 experts per token.
- The model has a 1M-token context window and native vision understanding.
- vLLM highlights Kimi Delta Attention, a hybrid attention design intended to make million-token context affordable.
- The team also thanked Moonshot, inferact, Nvidia, AMD, and the vLLM community for the launch integration.
Related event: vLLM Provides Day-0 Support for 2.8T Parameter Kimi K3(2 posts)→
More from Infra
- Gemma 4 is benchmarked locally on a 48GB Mac with MLX, llama.cpp and Java 25 — rseroter · 2026-07-28
- Bittensor subnet expansion is pitched as a cheaper AI infrastructure path for companies — markjeffrey · 2026-07-28
- Project Orion is training a 16B model live across three continents on heterogeneous compute — markjeffrey · 2026-07-28
- Modular handbook maps the hidden costs of LLM inference, from KV cache to prefill/decode splits — udmrzn · 2026-07-28
- Kimi K3 reaches Merge Gateway with U.S. inference providers and ZDR terms — shensi · 2026-07-28
- Compute, not algorithms, is the real moat in frontier AI — GavinSBaker · 2026-07-28