Hand-crafted Hessian diagonal estimator derivations: novel LLM training data?
Aiden_Tech_Ai · x · 2026-09-03
zhanpengzhou published a new blog: hand-crafted (no AI help) derivations of two estimators for the diagonal of the Hessian matrix, following the Sophia paper — a scalable second-order optimizer using a lightweight diagonal Hessian estimate as preconditioner, achieving 2x speedup over Adam on 125M–1.5B GPT models (same perplexity with 50% fewer steps). The author muses whether such purely hand-written derivations could serve as novel training data for LLMs.
More from Research
- Kirin builds large-scale animal motion dataset from in-the-wild video for 3D animation — Brian Nlong Zhao · 2026-09-03
- KAIST's Declarative Attention lets LLMs skip most KV cache reads — kaist-ai · 2026-09-03
- NVIDIA post-training pipeline hits gold-medal IOI performance, topping top humans — nvidia · 2026-09-03
- BPCO paper distills a stable PPO recipe for LLM RL, beating GRPO across scales — max_paperclips · 2026-09-03
- Puffin-World: open-source unified multimodal world model with native 3D states — ccloy · 2026-09-03
- LeVJEPA: video encoder matches V-JEPA 2 with 5.6-20.8x less pretraining compute — CSProfKGD · 2026-09-03