30 years on, deep learning still lacks a clear answer on flat minima
fleetwood___ · x · 2026-07-29
A post marking 30 years since Schmidhuber’s flat-minima paper says the field still doesn’t know whether flat minima generalize better.
The attached figure contrasts a “flat” and a “sharp” minimum, highlighting a long-running optimization question in deep learning: whether the geometry of the loss basin actually predicts generalization performance.
More from Research
- DLBCN 2026 opens presenter, spotlight and poster calls for Barcelona deep learning research — serrjoa · 2026-07-29
- NVIDIA’s PDD speeds image and video generation with 4–8-step distillation — nvidia · 2026-07-29
- TD-JEPA learns temporal progress from offline logs and beats LeWM on control tasks — HKBU-KnowComp · 2026-07-29
- NeurIPS timing shift may ease review pressure but could reshape ICLR submissions — deepakns · 2026-07-29
- Netherite rewrites Minecraft in C and CUDA, running 7,200 live worlds on one GPU — karmay007 · 2026-07-29
- Core Automation founders say transformer limits, not scale, are now the bottleneck — Training Data (Sequoia) · 2026-07-29