Inside Kimi's Trillion-Parameter MoE and Co-located RL Infrastructure

stochasticchasm · x · 2026-07-28

A deep dive into Kimi's trillion-parameter model reveals a sparser MoE architecture (16/896 activated parameters), likely bottlenecked by VRAM. It also highlights their rare, co-located reinforcement learning (RL) system at this massive scale and highly engineered sandbox infrastructure.

Related event: Kimi K3 Architecture: 3T Parameters and Ultra-Sparse MoE(2 posts)→

Original post →

More from Infra

Infra channel →