New paper finds 'anti-grokking': test accuracy collapses back to chance after successful generalization

sytelus · x · 2026-09-20

A new arXiv paper by Hari Prakash and Charles Martin extends two canonical grokking experiments (a 3-layer MLP on MNIST subset and a transformer on modular addition) far beyond standard training and reports a previously unreported third phase: anti-grokking — after transitioning to successful generalization, test accuracy collapses back to chance while training accuracy stays perfect.

Related event: Paper Reveals Third Stage of Grokking: Late-Stage Generalization Collapse(2 posts)→

Original post →

More from Research

Research channel →