Taylor Approximation Tracks Pre-training Loss

ShamKakade6 · x · 2026-07-09

Sham Kakade previewed an ICML 2026 talk discussing how far Taylor's theorem can go in deep learning, examining how accurate local quadratic approximations remain during the pre-training of large models. They found that on a 150 million parameter model, this approximation can track the true loss curve quite deep into training, suggesting that the local structure of the loss landscape might be more stable than intuitively expected.

Related event: Taylor Approximation Tracks LLM Pre-training Loss(2 posts)→

Original post →

More from Research

Research channel →