LLMs Memorize Then Generalize, Diffusion Models Generalize Then Memorize

gabriberton · x · 2026-07-03

Researchers point out that LLMs and visual diffusion models exhibit diametrically opposed learning patterns during training: LLMs tend to memorize training data first before gradually generalizing, whereas visual diffusion models demonstrate generalization capabilities first, with memorization occurring much later. A shared characteristic between the two model types is that increasing data volume is the most effective way to prevent excessive memorization. This comparison provides a fresh perspective on understanding the differences in overfitting and generalization mechanisms across varying architectures.

Original post →

More from Research

Research channel →