vLLM posts day-zero Kimi K3 support with 2.8T MoE serving details
vllm_project · x · 2026-07-29
The vLLM team and Inferact published a detailed guide on day-0 support for Kimi K3.
- They describe Kimi K3 as a 2.8T-parameter MoE model with 1M context and native vision.
- The post focuses on how vLLM handles KDA, MXFP4 MoE, KV cache management, prefill/decode disaggregation, and speculative decoding together.
- A ready-to-run command is provided, and the team says the easiest setup is 8 NVIDIA B300 or 8 AMD MI355X GPUs.
- Inferact also open-sourced a DSpark speculator for Kimi K3, with a concrete --speculative-config example.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23
- Cloudflare CTO Dane Knecht makes TIME's 2026 executives list as AI crawlers hit 52% of traffic — dinasaur_404 · 2026-09-23