SoftServe: Quasi-Newton Optimization Scales to Massive Neural Networks
SoftServe introduces scalable quasi-Newton methods for deep learning, handling non-convex objectives and massive networks on GPUs, outperforming Adam and Muon on ill-conditioned scientific tasks.
2026-10-02 ~ 2026-10-02 · 3 related posts
- SoftServe: A quasi-Newton method for non-convex objectives that scales to very large neural networks — dianarycai · 2026-10-02
- SoftServe preprint brings scalable quasi-Newton optimization to deep learning, beating Adam, Muon and SOAP on ill-conditioned tasks — dianarycai · 2026-10-02
- SoftServe's updates build on quadratic matrix equations, extending the authors' Batch-and-Match BBVI work — dianarycai · 2026-10-02