Scaling Tricks for 35B-A6B MoE
pbaylies · x · 2026-07-14
This is a progress report on training the 35B-A6B MoE model. The key finding is that the team successfully mitigated performance degradation when increasing the number of experts, without requiring additional FFN training. A specific insight shared: when scaling top-k from 8 to 32, they halved the output weights of the 9th to 32nd experts. The author notes that full details are available in the repository.
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22