Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

burny_tech · x · 2026-08-14

A new paper introduces the 'Information Abundance Paradox', revealing that during long-context pre-training, abundant relevant information in the context can reduce the model's incentive to encode that information parametrically, shifting its reliance toward contextualization.

The study shows that increasing the context window improves performance only up to an intermediate optimum, after which it consistently declines. During fine-tuning, richer train-time context improves results with supporting context but reduces robustness when context is absent or misleading at test time. Mechanistically, informative context shifts gradient pressure from feed-forward networks (linked to parametric knowledge) toward attention modules, as it provides a lower complexity solution.

Original post →

More from Research

Research channel →