LLM Memorization: The 3.6 Bits Per Parameter Threshold

gabriberton · x · 2026-07-03

The paper "How much do language models memorize?" reveals that LLMs continuously memorize training data until saturation occurs. This saturation point is reached at 3.6 bits per parameter; beyond this, the model begins to compress → generalize → exhibit double descent → and experience grokking. This establishes a quantitative threshold for the transition from memorization to generalization.

Original post →

More from Research

Research channel →