Pretraining study: varied auxiliary views beat repetition for LLM knowledge acquisition
kastnerkyle · x · 2026-09-10
A controlled pretraining study shows LLMs acquire knowledge better when the same content is presented in varied forms—textbook, Q&A, blog—dubbed "auxiliary views," rather than by repeating the source text.
- 36 documents from arXiv papers, US case law, and medical reports were fine-tuned into OLMo-2 at 1B/7B/13B/32B scales, with Qwen 7B as a cross-check; token counts were equalized across conditions
- Baselines: repeating source text vs. 9 paraphrases vs. paraphrases plus textbook/blog/Stack Exchange-style Q&A
- Notably, factual recall improved too—even though eval questions came straight from the source, auxiliary views beat plain repetition at retrieving those facts
- This suggests conceptual understanding aids memorization and generalized representations help recall of specific facts
- Paraphrases also mitigate overfitting from document repetition at small batch sizes; the views differ in explanation style, emphasized relations, assumed audience, and knowledge organization
More from Research
- Dev dives into hash-grid vs sparse-tree voxel simulation: latency is the bottleneck — Michael_Moroz_ · 2026-09-10
- CADArena benchmark shows AI builds accurate CAD geometry but unusable feature trees — hudzah · 2026-09-10
- AI Hasn't Solved Navier–Stokes: Mathematicians Led, LLMs Just Extended and Verified — PolarBearby · 2026-09-10
- Deep dive asks whether AlphaGenome Atlas's precomputed 9bn-variant lookup is a real unlock — Babayaga1664 · 2026-09-10
- Gensyn builds IR3DE-AXL, a decentralized collective inference network with no central gateway — benfielding · 2026-09-10
- If AI solves a Millennium Prize problem using human research, who gets credit? — Egologic · 2026-09-10