Study: LLM reasoning runs on procedural knowledge learned in pretraining, not memorized answers

egrefen · x · 2026-09-30

A paper by Laura Ruis et al. analyzes which pretraining documents drive LLM outputs on reasoning versus factual tasks. Using influence analysis on 7B and 35B models with 2.5B pretraining tokens each, they find factual questions rely on mostly distinct document sets, while different reasoning questions within a task share influential documents — evidence of procedural knowledge. Conclusion: LLM reasoning reflects reusable procedures learned during pretraining rather than parametrically retrieved answers, offering a new way to study generalization without train-test separation. The paper was accepted as an oral.

Related event: Study: LLM Reasoning Relies on Reusable Procedures Learned in Pretraining(2 posts)→

Original post →

More from Models

Models channel →