DeepSeek vs Kimi MoE Config Differences
teortaxesTex · x · 2026-07-17
The author updated their assessment of MoE routing scales:
- Initially, it was expected that DeepSeek would adopt a 16 active experts / large expert pool route.
- However, the actual implementation is 7/384, while Kimi uses 16/896.
- This indicates that whether to continue scaling "expert granularity" isn't a straightforward conclusion, and the industry still explores different implementation paths.
The main point isn't to praise a specific model, but to highlight that MoE expert granularity design is still an active exploration area, with vendors adopting different strategies for active expert counts and total pool sizes.
More from Models
- Grok 4.5 is now free inside Cursor, the popular AI coding IDE — mark_k · 2026-07-21
- GPT often converges on the same near-miss ideas in math problems — yacineMTB · 2026-07-21
- Eno Reyes says model distillation is basically unstoppable — LangChain · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- OpenAI hackathon project stalls as Codex struggles on voice, while Claude spots the issue — ColleenMBrady · 2026-07-21
- Kimi K3 lands exactly on China’s 2-year AI capability trend line — peterwildeford · 2026-07-21