Quadratic Models Surprisingly Accurately Describe LLM Pretraining, Paper Finds
jasondeanlee · x · 2026-07-30
Researcher alexmeterez shared findings on how well optimization theory's simplest model—the quadratic—describes LLM pretraining, showing it works surprisingly well.
By linearizing a 150M parameter model at checkpoints during training, the local Taylor expansion accurately tracks the true loss for up to 10% of the training budget.
More from Research
- UCSD Paper Introduces LeRoPE: A Superior and Efficient Upgrade to Rotary Positional Encodings — burkov · 2026-07-30
- ICML 2026 Paper Explorer Launched with 6,341 Papers — algo_diver · 2026-07-30
- Compiling Fuzzy Functions Directly into Neural Weights: The ProgramAsWeights Paradigm — weichiuma · 2026-07-30
- EMBC 2026: Predicting Multiple Neuropathologies with Interpretable AutoML — PTenigma · 2026-07-30
- Reddit Debate: Is ARC-AGI 3 an Intentionally Dishonest Measure of AGI? — Glittering-Neck-2505 · 2026-07-30
- When Does Synthetic Data Work? Research Reveals Optimal Ratios and 'Zeta Law' — PTenigma · 2026-07-30