Long-context training may undermine parametric knowledge: Info Abundance Paradox
DanielKhashabi · x · 2026-08-18
The paper proposes the "Information Abundance Paradox," suggesting that abundant task-relevant context reduces the need to store knowledge in model parameters. Performance improves only up to an intermediate training context length before robustness declines. As context grows, optimization shifts from FFNs toward attention, altering how the model represents knowledge.
More from Research
- NVIDIA Open-Sources VoiceChat, a Full-Duplex Speech-to-Speech Model with Tool Calling — chaumian · 2026-09-22
- SteerDuplex: full-duplex speech model gains 44.5pp in steerability, new SteerBench released — ScaleAI · 2026-09-22
- Sarah Hooker suspects fake AI paper submissions cluster at a few universities — sarahookr · 2026-09-22
- Rollout scheduling: the underappreciated infra trick boosting inference efficiency — stochasticchasm · 2026-09-22
- Benchling benchmarks Claude and ChatGPT on wet-lab protocols: helpful, not solved — nlarusstone · 2026-09-22
- LLMs as probability-guided search in token space: data coverage gaps are the real weakness — AlexTensor · 2026-09-22