Do we still need LR warmup for transformer finetunes? Maybe as a spiky WSD variant

kalomaze · x · 2026-08-19

Someone asked whether, in modern transformer finetuning (general SFT), LR warmup still reliably improves evals — or whether the need has been obviated.

kalomaze's take: it might make sense if you want a briefly higher peak LR near the start — like a spiky warmup modification of WSD scheduling, rather than the traditional slow ramp.

Related event: Is LR Warmup Still Needed in Modern Transformer Fine-Tuning?(2 posts)→

Original post →

More from Research

Research channel →