Quadratic Models Accurately Predict LLM Pretraining Loss, Study Shows

ShamKakade6 · x · 2026-07-30

A recent study reveals that the simplest model in optimization theory, the quadratic, accurately describes LLM pretraining. By linearizing a 150M parameter model at various training checkpoints, researchers found that the local Taylor expansion tracks the true loss effectively for up to 10% of the training budget. The paper is predicted to become highly influential in the AI community.

Related event: Quadratic Models Can Accurately Describe LLM Pre-training Dynamics(3 posts)→

Original post →

More from Research

Research channel →