Scaling Tricks for 35B-A6B MoE

pbaylies · x · 2026-07-14

This is a progress report on training the 35B-A6B MoE model. The key finding is that the team successfully mitigated performance degradation when increasing the number of experts, without requiring additional FFN training. A specific insight shared: when scaling top-k from 8 to 32, they halved the output weights of the 9th to 32nd experts. The author notes that full details are available in the repository.

Original post →

More from Research

Research channel →