A new paper argues flat minima may be an illusion in deep learning

theomitsa · x · 2026-07-28

A new paper asks whether flat minima are an illusion.

The author argues that flatness in parameter space is not the same thing as functional simplicity. By rescaling ReLU networks, the paper shows the raw Hessian trace can change by up to 99× while predictions stay fixed. It proposes a reparameterization-invariant notion called weakness and reports that a joint completion score correlates with held-out accuracy, while raw curvature metrics do not.

Original post →

More from Research

Research channel →