SoftServe preprint brings scalable quasi-Newton optimization to deep learning, beating Adam, Muon and SOAP on ill-conditioned tasks
dianarycai · x · 2026-10-02
Researchers including Diana Cai released SoftServe, a family of quasi-Newton optimization methods designed for non-convex deep learning objectives that scales to very large networks.
- Derives positive-definite curvature estimates from a variational objective, with no line searches or ad hoc curvature corrections.
- Diagonal and Kronecker-factored variants preserve positive definiteness by construction and scale to massive models.
- Uses a stable coupled Newton-Schulz iteration, replacing costly matrix decompositions with GPU-friendly matrix multiplications.
- Excels on severely ill-conditioned problems — recurrent networks, deep autoencoders, physics-informed neural networks and a 136M-parameter physics-informed diffusion model — often achieving lower losses than Adam, Muon and SOAP.
Preprint: arXiv:2610.02182.
Related event: SoftServe: Quasi-Newton Optimization Scales to Massive Neural Networks(3 posts)→
More from Research
- AI 'speech clock' predicts how fast you're ageing from your voice, featured in Nature — AnnaCiaunica · 2026-10-02
- ScholarCatalyst: a new benchmark testing whether AI can pick research problems like humans — PangWeiKoh · 2026-10-02
- Clarification as Supervision lands NeurIPS Oral: denser training signal via model interaction — iatitov · 2026-10-02
- Ego-Exo4D-HM: SMPL-H reconstructions for 523 hours of egocentric-exocentric video, open-sourced — geopavlakos · 2026-10-02
- Cohere Labs Talk: Making AI Math Reasoning Machine-Checkable with Lean — Cohere_Labs · 2026-10-02
- New paper dissects only task-relevant network parts, making mechanisms inspectable and editable at far lower cost — leedsharkey · 2026-10-02