SoftServe: GPU-scale quasi-Newton that often beats Adam and Muon on ill-conditioned tasks

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower

cs.LG, cs.AI

2026-10-02

SoftServe softens the BFGS secant constraint and uses diagonal/Kronecker curvature, scaling quasi-Newton to 136M models and often beating Adam and Muon on ill-conditioned tasks.

What problem this solves

Quasi-Newton (QN) methods such as BFGS dominate large-scale convex optimization, yet deep learning trains almost exclusively with first-order methods. Two obstacles explain the gap. Non-convexity: the loss Hessian has negative eigenvalues, and an inverse-Hessian estimate that inherits one can turn the update into an ascent direction. BFGS keeps its estimate positive definite through the curvature condition sᵀy>0, which non-convex objectives violate; enforcing the secant equation anyway forces the estimate to become indefinite. The textbook fix, a line search, needs full-batch losses and does not fit stochastic minibatch training. Scale: a dense estimate stores D² entries, and L-BFGS with ten difference pairs effectively stores twenty copies of the model state.

Both obstacles bite hardest exactly where curvature would help most. PINNs and recurrent networks have badly conditioned losses on which first-order methods crawl.

Method

SoftServe builds on Soft QN (Berglund et al., 2025), which replaces the secant equation Hy=s with a least-squares penalty regularized by a log-det divergence toward the previous estimate. The objective always has a unique positive-definite solution, even under negative curvature, and recovers BFGS in the limit of positive curvature with λ→∞. The catch is cost: the dense solution needs O(D²) time and memory. SoftServe instead solves the same objective over structured families:

These equations are solved without SVD, whose common implementations suit GPUs poorly. Matrix square roots and inverses instead come from the coupled Newton–Schulz iteration, built from matrix multiplications and additions only (18 iterations for the square root, 10 for the inverse in the implementation). This is the same numerical machinery behind Muon's scalability. Three practical pieces complete the recipe: momentum with Nesterov extrapolation or bias correction; a normalized preconditioned step, meaning descent constrained to a ball in the H-metric; and a curvature refresh every K=10 steps. At each window start the optimizer snapshots the parameters, the gradient, and every random choice (minibatch indices, dropout masks, augmentation); at window end it re-evaluates the gradient under the same draw, producing one clean secant pair whose difference reflects the curvature of a single sampled loss. The replay costs about one extra gradient evaluation per ten steps, roughly 10% overhead, and every comparison charges it against SoftServe's budget.

Results

Baselines: Adam, Muon, SOAP, K-FAC, and K-BFGS(L), plus L-BFGS on deterministic tasks, plus SoftServe-I (H fixed to the identity) as an ablation. All methods get equal gradient budgets, per-method learning-rate sweeps, and three fresh seeds.

TaskMetricResult
RNN Adding (321 params)test MSESoftServe-Kron 6.6e-6; Muon 1.5e-4; SoftServe-I 1.0e-2
MNIST autoencoder (2.8M)training objectivelowest in full batch; close to Muon at batch size 1,000
Small PINNs (Wave/Convection/Reaction, fixed or resampled points)final training objectivelowest in 5 of 6 settings; fixed-point Wave is the exception, where Adam+L-BFGS wins
PirateNet (0.73M)mean losslowest on KdV; on par with Muon on Allen–Cahn
Physics-informed diffusion (PIDM, 136M)training loss at 100k gradient evaluationsbelow Adam, Muon, and SOAP
16M GPT (200M FineWeb tokens)validation NLLSoftServe-Kron 4.661; AdamW 4.786; Muon 4.440; SOAP 4.416

On RNN Adding the full loss Hessian is numerically near-singular at initialization across all four seeds. Preconditioning is the whole story there: without curvature, SoftServe-I stalls near 1e-2; with Kronecker curvature it reaches the 1e-6 range.

The language model is an explicit negative result: SoftServe beats AdamW and loses to Muon and SOAP.

Why it matters

For AI4Science practitioners this is a tool to try today. PINNs, PDE solvers, and physics-coupled generative models have exactly the ill-conditioned losses where curvature pays, and the code is public. For optimizer researchers, the paper shows a clean division of labor: positive definiteness guaranteed by construction through the variational objective, with no damping or line searches, and scalability from Kronecker structure plus Newton–Schulz, the same GPU building blocks Muon uses. It is not a general-purpose optimizer: the language-model result shows no advantage in LM pretraining, and 136M is the largest model tested. The honest positioning is a solid systems step that brings quasi-Newton training to model sizes it could previously not touch, with clearly drawn boundaries.

Limitations

The authors' own list: λ is fixed within each run, with adaptive λ left to future work, and convergence guarantees under realistic settings remain open. Reading the experiments adds more. The wins concentrate on ill-conditioned scientific tasks; on the LM the method trails Muon and SOAP. The exact Kron refresh is cubic in the layer dimensions and is amortized only by refreshing every ten steps, with wall-clock comparisons confined to the appendix. λ needs per-task tuning and selected values span 1 to 9999, four orders of magnitude, so moving to a new task is not free. The replay mechanism requires storing and replaying every source of randomness, an engineering constraint most optimizers do not impose.

Terms

Source

What people are saying

Related papers

All paper explainers