Information Abundance Paradox: Long-context training undermines parametric knowledge

DanielKhashabi · x · 2026-08-16

Researchers at Johns Hopkins University propose the Information Abundance Paradox: abundant relevant information in training context reduces the incentive to encode it parametrically, increasing reliance on context. In pretraining, longer contexts improve language modeling, NLU, and closed-book MCQA only up to an intermediate optimum, then decline. In SFT, more context improves performance with support but reduces robustness when context is absent or misleading. Mechanistically, training with informative context shifts gradient pressure from feed-forward networks.

Related event: Information Abundance Paradox: Long-Context Training Can Hurt Parametric Knowledge and Short Tasks(5 posts)→

Original post →

More from Models

Models channel →