Randomizing SWA spans for 20% of samples improves full-seqlen loss, finds training experiment
kalomaze · x · 2026-10-07
kalomaze reports a toy baseline experiment: randomizing the power-of-2 span from 128 to 512 for full SWA on all layers for 20% of samples showed obvious improvements — speeding up full-seqlen loss with less total information, unlike normal data augmentation. The underlying principle: even RLVR with held-out answers forces the model to derive correct answers via reasoning, a deliberately manufactured, far more brutal information asymmetry.
More from Research
- NeuroAI manifesto: applying AI scaling laws to BCIs, the bitter lesson for the brain — w1kke · 2026-10-07
- OpenAI releases 722 math papers from an internal frontier model, including a claimed quasi-Riemann proof — imjustnewatai · 2026-10-07
- LeanLean benchmark: Opus 5.5 scores 64.3% compressing Lean proofs, GPT 6.1 Sol only 39.9% — ChrSzegedy · 2026-10-07
- PersistBench (NeurIPS Spotlight): 4D foundation models can see but not remember — weichiuma · 2026-10-07
- COLM 2026 poster: Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer — boknilev · 2026-10-07
- AI-Read Gold Electrodes Detect Molecular Chirality One Molecule at a Time — Brighter-Side-News · 2026-10-07