Taylor Approximation Tracks Pre-training Loss
ShamKakade6 · x · 2026-07-09
Sham Kakade previewed an ICML 2026 talk discussing how far Taylor's theorem can go in deep learning, examining how accurate local quadratic approximations remain during the pre-training of large models. They found that on a 150 million parameter model, this approximation can track the true loss curve quite deep into training, suggesting that the local structure of the loss landscape might be more stable than intuitively expected.
Related event: Taylor Approximation Tracks LLM Pre-training Loss(2 posts)→
More from Research
- OpenForecaster uses daily news to improve language-model forecasting — Cohere_Labs · 2026-07-21
- SenseTime unveils U1 Pro and open-sources a 50M-sample vision dataset at WAIC 2026 — 机器之心 · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21