Kimi K3 pretraining combines PP, EP and ZeRO-1 with balanced MoE routing

SeunghyunSEO7 · x · 2026-07-28

Kimi K3’s pretraining stack mixes PP, EP, ZeRO-1 and context parallelism

The attached technical report page says Kimi K3 pretraining combines pipeline parallelism (with virtual stages), expert parallelism, ZeRO-1 data parallelism, pipeline ZeRO-2 gradient sharing, and context parallelism for 3T-class pretraining.

The paper section on MoonEP explains the motivation:

The key claim is that the system keeps the overall computation flow of conventional MoE training while making load balance and communication much cleaner at scale.

Related event: Kimi K3 Tech Report: 2.8T MoE and Architectural Efficiency Breakthroughs(49 posts)→

Original post →

More from Infra

Infra channel →