Transformers lose length extrapolation as training progresses, weight decay speeds the decay

andre_t_martins · x · 2026-10-08

New research from Pavlo Vasylenko and collaborators at Sardine Lab (with Marcos Treviso and Matthias Lindemann) examines why neural models fail to generalize to longer sequences.

The findings suggest regularization choices affect length extrapolation specifically, with direct implications for long-context reasoning workloads.

Original post →

More from Research

Research channel →