Study: LLM reasoning runs on procedural knowledge learned in pretraining, not memorized answers
egrefen · x · 2026-09-30
A paper by Laura Ruis et al. analyzes which pretraining documents drive LLM outputs on reasoning versus factual tasks. Using influence analysis on 7B and 35B models with 2.5B pretraining tokens each, they find factual questions rely on mostly distinct document sets, while different reasoning questions within a task share influential documents — evidence of procedural knowledge. Conclusion: LLM reasoning reflects reusable procedures learned during pretraining rather than parametrically retrieved answers, offering a new way to study generalization without train-test separation. The paper was accepted as an oral.
Related event: Study: LLM Reasoning Relies on Reusable Procedures Learned in Pretraining(2 posts)→
More from Models
- Opus 5.5 builds a working computer from scratch: 277k logic gates, OS and games in JS — LeviTurk · 2026-09-30
- Insider claim: OpenAI staff only test products via Slack, while Astra criticized as too slow — zephyr_z9 · 2026-09-30
- OpenAI's dots is slow to answer and respond, likely overloaded at launch — tinyfool · 2026-09-30
- Benchmarking Qwen 3.8-Flash-Next on Strix Halo: Halogen hits 1,045 prefill t/s — deepu105 · 2026-09-30
- Why hasn't Gemini 4.0 dropped? Reddit community speculates on Google's delayed flagship — Jumpy-Cobbler1020 · 2026-09-30
- User A/B test suggests Opus 5.5 output quality shifted noticeably within a week — skelzer · 2026-09-30