Moonshot’s Mooncake serving stack boosts long-context throughput by up to 525%

stochasticchasm · x · 2026-07-28

Mooncake is a KV cache–centric serving architecture for Kimi, Moonshot AI’s LLM service.

The paper describes a disaggregated design that separates prefill and decoding clusters and uses underutilized CPU, DRAM, and SSD resources in the GPU cluster as an external KV cache pool.

Main ideas:

Reported results:

The paper is especially focused on long-context serving and long-context RL rollouts.

Original post →

More from Infra

Infra channel →