Transformers lose length extrapolation as training progresses, weight decay speeds the decay
andre_t_martins · x · 2026-10-08
New research from Pavlo Vasylenko and collaborators at Sardine Lab (with Marcos Treviso and Matthias Lindemann) examines why neural models fail to generalize to longer sequences.
- It's not just an architectural problem: transformers can lose their extrapolation ability as training progresses
- Weight decay accelerates this degradation
- Properly placed dropout improves extrapolation
The findings suggest regularization choices affect length extrapolation specifically, with direct implications for long-context reasoning workloads.
More from Research
- OpenAI preprint claims matrix multiplication exponent drops to 2.25, biggest leap since 1979 — BorisMPower · 2026-10-08
- Google's new federated-learning design logs server access policies publicly — Crescitaly · 2026-10-08
- AI2's Bolmo tackles the 'token tax' hitting Global South scripts — Kyle_L_Wiggers · 2026-10-08
- TEMPO adds temporal context to VLAs, lifting robot bottle handover success from 44% to 74% — _krishna_murthy · 2026-10-08
- CheckerBench: Best Coding Agent Scores Just 45.33% on Static-Analysis Checker Synthesis — humanlaya-data-lab · 2026-10-08
- Gary Marcus: OpenAI's vague math report 'would never pass peer review' — Tao responds too — Gary Marcus · 2026-10-08