LLM Pretraining Study: Factual Knowledge Peaks ~30 Steps After Exposure, Duplicated Data Speeds Forgetting
gordic_aleksa · x · 2026-09-10
An arXiv study by Hoyeon Chang, Minjoon Seo et al. dissects how LLMs acquire factual knowledge during pretraining.
- Delayed memorization: after a fact is injected at step t=0, peak memorization occurs 30 steps later — even though subsequent batches don't contain the fact. Acquisition works by gradually raising the fact's output probability, then being diluted by forgetting.
- More data doesn't help: counterintuitively, pretraining on more data shows no significant improvement in acquiring and retaining factual knowledge.
- Power-law forgetting: memorization and generalization both decay as a power law with training steps, and duplicated training data speeds up forgetting.
- Larger batches help robustness against forgetting.
These findings offer plausible explanations for observed LLM behaviors like poor knowledge updating, with practical implications for data curation and knowledge injection.
More from Research
- GPN-Star: phylogeny-aware genomic language model hits SOTA on variant effect prediction — pastramimachine · 2026-09-10
- Indie researcher releases open-source audio model that turns text prompts into playable synths — RoyalCities · 2026-09-10
- Grok: Clay Institute prize is far off — OpenAI's math result still needs extensive vetting — MikePFrank · 2026-09-10
- Artificial Analysis launches Optima to build custom benchmarks, testing GPT-6 Astra — ArtificialAnlys · 2026-09-10
- TUM releases NOAH, a longitudinal multimodal time-aware model for full patient journeys — TUM-AIMED · 2026-09-10
- Will Automating AI R&D Trigger a Software Intelligence Explosion? Paper Analyzes — nabeelqu · 2026-09-10