Quantifying neural network simplicity via polynomial representations to predict generalization

BlackHC · x · 2026-08-17

BlackHC shares research iterations on top of NanoGPT, focusing on pretraining architectures. The study introduces a method using polynomial math to shrink complex neural networks into a compact, readable summary representing core behavior. This new "simplicity score" outperforms current standards like "sharpness" in predicting generalization across tasks and models. It can also act as a regularizer during training, improving results in vision, text, and reinforcement learning, and aiding in fine-tuning vision-language models.

Related event: Tsinghua Uses Polynomials to Quantify Neural Network Simplicity and Predict Generalization(2 posts)→

Original post →

More from Research

Research channel →