vLLM lists Kimi K3 serving options with disaggregation, KV offload and MoE backends
vllm_project · x · 2026-07-27
- The post lists the main serving features that can be enabled for Kimi K3 on vLLM.
- Highlights include prefill/decode disaggregation, topology-aware MoE backends (megamoe for EP, trtllm for TP>1), KV offloading, and support for tool calling, reasoning, and structured output.
- The attached screenshot shows the vLLM Recipes page for moonshotai/Kimi-K3, including multi-node tensor parallel deployment, hardware targets, and latency-oriented serving settings.
More from Infra
- Claude chat indexing incident sparks a push for confidential-compute AI — bittingthembits · 2026-07-27
- Moonshot open-sources MoonEP as open models vs closed labs debate intensifies — KyeGomezB · 2026-07-27
- Kimi K3’s 2.5x scaling-law gain draws praise for training efficiency — andrew_n_carr · 2026-07-27
- Kimi K3 goes live on Nebius with 1M-token context and a 57 AA score — teortaxesTex · 2026-07-27
- AI agent finds a longstanding Bun Node-compat bug in `child_process.spawn` — steipete · 2026-07-27
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27