Study Reveals Richness Paradox: Long-Context Training Weakens Parametric Memory

A new study identifies an "information richness paradox": when training contexts are rich with relevant information, models lose the incentive to internalize knowledge into parameters. Performance improves with longer context only up to a sweet spot, after which parametric learning degrades, verified at 10B-token scale.

2026-08-16 ~ 2026-08-18 · 2 related posts