Moonshot’s Mooncake serving stack boosts long-context throughput by up to 525%
stochasticchasm · x · 2026-07-28
Mooncake is a KV cache–centric serving architecture for Kimi, Moonshot AI’s LLM service.
The paper describes a disaggregated design that separates prefill and decoding clusters and uses underutilized CPU, DRAM, and SSD resources in the GPU cluster as an external KV cache pool.
Main ideas:
- keep active decoding blocks in GPU KV cache
- write reusable idle prefixes back to CPU DRAM only after eviction
- prefetch them before reuse
- offload training states to NVMe after each training iteration to free DRAM
- add a prediction-based early rejection policy for overloaded rollout scenarios
Reported results:
- up to 525% throughput improvement in some simulated long-context scenarios
- 75% more requests handled under real workloads
The paper is especially focused on long-context serving and long-context RL rollouts.
More from Infra
- Optimizing K8s Resource Requests Yields 9x Speedup for Whisper Workloads — anacondainc · 2026-07-28
- Linear-attention hybrids may need finer caching for long prompts and workflows — stochasticchasm · 2026-07-28
- A research-agent ranking of 8 stock-data MCP servers puts Equibles first — DanielAPO · 2026-07-28
- Underlayer Electrons Aggravate Stochastic Defectivity in EUV Lithography — CatAstro_Piyush · 2026-07-28
- Moonshot’s Kimi K3 report details a microVM sandbox system for agentic RL — stochasticchasm · 2026-07-28
- SkyPilot says serving Kimi K3 needs multi-node inference and a full stack — skypilot_org · 2026-07-28