Information Abundance Paradox: long-context training can hurt short-task performance

DanielKhashabi · x · 2026-08-15

NLP researcher Daniel Khashabi describes what he calls the "Information Abundance Paradox": once training context length passes a certain threshold, performance on short-context tasks actually drops.

His explanation: in long-context training, relevant knowledge is readily available in the context, so the model has less incentive to internalize that knowledge into its parameters — hurting it when the knowledge isn't right there in the prompt.

Related event: Study Proposes 'Information Abundance Paradox' in Long Context Training(4 posts)→

Original post →

More from Models

Models channel →