Preprint finds longer context can weaken parametric learning in LLMs
DanielKhashabi · x · 2026-08-14
A new preprint proposes the "Information Abundance Paradox": when relevant information is abundant in the training context, the model has less incentive to internalize that information into its parameters — meaning longer context can actually weaken parametric learning in LLMs.
In other words, the more ready-made answers sit in the context, the more the model leans on copying from context rather than writing knowledge into its weights, with direct implications for long-context training and data-mixing strategies.
Related event: Study Proposes 'Information Abundance Paradox' in Long Context Training(4 posts)→
More from Research
- GitHub Project Uses AI to Mass-Produce 0days to Force Vendor Fixes — evilsocket · 2026-08-15
- Dynamic Short Convolutions Improve Transformer Performance — Cohere · 2026-08-15
- Paper: ML tools could link hippocampal function theories to real neural computations — AndrewLampinen · 2026-08-15
- TEMPO: recursive self-critique plus macro-step value estimation, not exploration rewards — teortaxesTex · 2026-08-15
- Microsoft Proposes RadFusion for Threshold-Controllable Radiology Reports — erichorvitz · 2026-08-15
- New Paper Proposes 'Latent Reading' to Interpret Literature via Machine Learning — begusgasper · 2026-08-15