Linear Regression Can Also Experience Grokking

RichmanRonald · x · 2026-07-14

This referenced paper and thread discuss an interesting conclusion: **grokking is not unique to deep neural networks**. The author points out that classic grokking can be observed even in simple, linear, and over-parameterized ridge regression: - Perfect fitting is achieved on the training set very early on - But generalization performance suddenly improves dramatically much later This suggests that grokking might reflect more general optimization and generalization dynamics, rather than just being a quirk of specific neural networks.

Original post →

More from Research

Research channel →