Do we still need LR warmup for transformer finetunes? Maybe as a spiky WSD variant
kalomaze · x · 2026-08-19
Someone asked whether, in modern transformer finetuning (general SFT), LR warmup still reliably improves evals — or whether the need has been obviated.
kalomaze's take: it might make sense if you want a briefly higher peak LR near the start — like a spiky warmup modification of WSD scheduling, rather than the traditional slow ramp.
Related event: Is LR Warmup Still Needed in Modern Transformer Fine-Tuning?(2 posts)→
More from Research
- ACL 2026 Paper: LLMs Stick to Old Knowledge Despite New Information — mdredze · 2026-08-19
- NeurIPS 2026 Workshop Focus: AI Writing, AI Review, and Academic Governance — ManlingLi_ · 2026-08-19
- OpenAI Reply: RLHF and World Models Accelerate Design Iteration — TinfoilTricorn · 2026-08-19
- A self-devised 'freeze a service first' method let Codex pass a Terminal-Bench task with zero public solves — Present-Quantity-813 · 2026-08-19
- Stanford, Yale Win First Live AI Agent Competition at Data + AI Summit — CShorten30 · 2026-08-19
- Tuesday index combines 8 Surge benchmarks, more coming — echen · 2026-08-19