Paper: Long Context Weakens Parametric Learning in LLMs

SonglinYang4 · x · 2026-08-16

A new preprint proposes the "Information Abundance Paradox": when relevant information is abundant in the training context, the model has less incentive to internalize it into parameters. This weakens parametric memory and holds across model sizes trained on 10B tokens.

Related event: Study Reveals Richness Paradox: Long-Context Training Weakens Parametric Memory(2 posts)→

Original post →

More from Research

Research channel →