Kimi K3 is said to cost more than 2× as much to serve as V4
teortaxesTex · x · 2026-07-27
A reply thread around Kimi K3 argues that the model is far more expensive to serve than V4, reportedly by more than 2×.
The commenter also says Moonshot had to solve a large number of difficult engineering problems, and that the model report does not disclose training token counts. The main takeaway is that Kimi K3 may be a very large and capable frontier model, but its deployment economics look heavy.
More from Infra
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27
- Open-sourced Kimi K3 speculator lifts single-stream throughput from 118 to 370 tok/s — vllm_project · 2026-07-27
- vLLM lists Kimi K3 serving options with disaggregation, KV offload and MoE backends — vllm_project · 2026-07-27
- Inferact’s Kimi-K3-DSpark draft model reuses MLA caches to speed up vLLM serving — vllm_project · 2026-07-27
- Kimi K3 uses fixed-size KDA state instead of a growing KV cache — vllm_project · 2026-07-27
- MoonshotAI open-sources MoonEP for perfectly balanced expert parallelism — teortaxesTex · 2026-07-27