Karpathy's 'overfit first, regularize later' still rules large-scale post-training

rdesh26 · x · 2026-09-05

Karpathy's classic A Recipe for Training Neural Networks methodology still holds in the large-scale post-training era. A reader of a new long-form writeup highlighted its practical details:

The core principle — "overfit first, regularize later" — spans the whole evolution from the CNN era to RL post-training.

Original post →

More from coding & agent

coding & agent channel →