Quadratic Models Can Accurately Describe LLM Pre-training Dynamics

A recent study reveals that simple quadratic models can surprisingly and accurately describe the pre-training dynamics of Large Language Models. By linearizing checkpoints of a 150M parameter model, researchers found that local Taylor expansions closely track actual losses, particularly during the later stages of training.

2026-07-30 ~ 2026-07-31 · 3 related posts