Kimi K3’s MoE routing may be driving higher expert-parallel communication costs

stochasticchasm · x · 2026-07-28

Kimi K3’s MoE design may be paying extra communication cost for top-k 16 routing

The discussion points to the report’s Stable LatentMoE section. The key idea is that Kimi K3 uses a LatentMoE-style split between the full model width and the routed expert width, which lets it scale channel mixing to 896 routed experts with 16 active experts per token.

The poster argues that two things may be driving expensive expert-parallel communication: high sparsity and the choice of top-k 16 instead of top-k 8, since wider routing increases the chance that tokens must be sent to more nodes. The screenshot also notes a very aggressive 12-layer attnres setup, likely influenced by pipeline parallelism.

Original post →

More from Models

Models channel →