Hand-crafted Hessian diagonal estimator derivations: novel LLM training data?

Aiden_Tech_Ai · x · 2026-09-03

zhanpengzhou published a new blog: hand-crafted (no AI help) derivations of two estimators for the diagonal of the Hessian matrix, following the Sophia paper — a scalable second-order optimizer using a lightweight diagonal Hessian estimate as preconditioner, achieving 2x speedup over Adam on 125M–1.5B GPT models (same perplexity with 50% fewer steps). The author muses whether such purely hand-written derivations could serve as novel training data for LLMs.

Original post →

More from Research

Research channel →