Long-context training may undermine parametric knowledge: Info Abundance Paradox

DanielKhashabi · x · 2026-08-18

The paper proposes the "Information Abundance Paradox," suggesting that abundant task-relevant context reduces the need to store knowledge in model parameters. Performance improves only up to an intermediate training context length before robustness declines. As context grows, optimization shifts from FFNs toward attention, altering how the model represents knowledge.

Related event: Study Reveals Richness Paradox: Long-Context Training Weakens Parametric Memory(2 posts)→

Original post →

More from Research

Research channel →