Information Abundance Paradox: Long-context training undermines parametric knowledge
DanielKhashabi · x · 2026-08-16
Researchers at Johns Hopkins University propose the Information Abundance Paradox: abundant relevant information in training context reduces the incentive to encode it parametrically, increasing reliance on context. In pretraining, longer contexts improve language modeling, NLU, and closed-book MCQA only up to an intermediate optimum, then decline. In SFT, more context improves performance with support but reduces robustness when context is absent or misleading. Mechanistically, training with informative context shifts gradient pressure from feed-forward networks.
More from Models
- Anthropic details how Claude’s new watermarks work: mechanism and resistance to editing — TechCrunch AI · 2026-08-16
- Qwen3.8-27B on 16GB VRAM: KV Cache Quantization Cliff from q4_0 to q4_1 — Unnamed-3891 · 2026-08-16
- Benchmark: Qwen3.8-27B doubles coding ability vs Qwen3.6-27B — poppear · 2026-08-16
- Qwen 3.8 27B Beats Codex in Coding Benchmarks: Wins 8/13, Costs 1/3 — tokenbender · 2026-08-16
- DeepSeek Accused of Grey Testing for High Scores, Opus 5 Output Quality Questioned — teortaxesTex · 2026-08-16
- Study: Cross-Version Transfer of Qwen Interpretability Lenses — imstilllearningthis · 2026-08-16