Study Reveals Richness Paradox: Long-Context Training Weakens Parametric Memory
A new study identifies an "information richness paradox": when training contexts are rich with relevant information, models lose the incentive to internalize knowledge into parameters. Performance improves with longer context only up to a sweet spot, after which parametric learning degrades, verified at 10B-token scale.
2026-08-16 ~ 2026-08-18 · 2 related posts
- Paper: Long Context Weakens Parametric Learning in LLMs — SonglinYang4 · 2026-08-16
- Long-context training may undermine parametric knowledge: Info Abundance Paradox — DanielKhashabi · 2026-08-18