Degenerate Fisher information explains why huge neural nets don't defy Occam's razor
FrnkNlsn · x · 2026-09-10
A paper by Ke Sun and Frank Nielsen (arXiv:1905.11027, first posted 2019, revised to v9 in 2025) tackles a classic puzzle: deep networks have massive parameter spaces, seemingly defying Occam's razor, yet generalize well.
- Approach: define a locally varying dimensionality of the parameter space via the number of significant dimensions of the Fisher information matrix, modeling it as a singular semi-Riemannian manifold.
- Finding: DNN Fisher information matrices are nearly degenerate, so effective complexity is far lower than raw parameter count, yielding short description lengths under an MDL framework.
- Significance: provides a geometric information-theoretic bound explaining why large networks remain "simple" in the sense that matters for generalization.
More from Research
- Deep dive asks whether AlphaGenome Atlas's precomputed 9bn-variant lookup is a real unlock — Babayaga1664 · 2026-09-10
- Gensyn builds IR3DE-AXL, a decentralized collective inference network with no central gateway — benfielding · 2026-09-10
- If AI solves a Millennium Prize problem using human research, who gets credit? — Egologic · 2026-09-10
- Pretraining study: varied auxiliary views beat repetition for LLM knowledge acquisition — kastnerkyle · 2026-09-10
- ByteDance Seed Unveils ByteWrist: A Parallel Robotic Wrist for Confined-Space Manipulation — scott_e_reed · 2026-09-10
- The Prism Hypothesis unified autoencoding paper accepted at ECCV 2026, paves way for encoder-free MLLMs — liuziwei7 · 2026-09-10