Token Superposition Training Could Cut MoE Pretraining Compute by 2.5x

Aleksa Gordić explains Nous Research's token superposition training (TST), which processes overlapping groups of tokens early in pretraining and reportedly achieves the same loss with a 10B MoE model at 2.5x less compute, though one researcher reports his similar experiments yielded poor results.

2026-09-19 ~ 2026-09-19 · 2 related posts