Randomizing SWA spans for 20% of samples improves full-seqlen loss, finds training experiment

kalomaze · x · 2026-10-07

kalomaze reports a toy baseline experiment: randomizing the power-of-2 span from 128 to 512 for full SWA on all layers for 20% of samples showed obvious improvements — speeding up full-seqlen loss with less total information, unlike normal data augmentation. The underlying principle: even RLVR with held-out answers forces the model to derive correct answers via reasoning, a deliberately manufactured, far more brutal information asymmetry.

Original post →

More from Research

Research channel →