Georgia Tech Traces LLM Reasoning to Training Data via Open OLMo
allen_ai · x · 2026-08-22
A Georgia Tech team used the fully open OLMo ecosystem (model, Dolma dataset, WebOrganizer) and influence functions to trace model performance in social reasoning and knowledge back to specific training text types. They found dialogue-rich, interpersonal writing significantly impacts reasoning more than factual knowledge. This research was only possible with a fully open model stack.
Related event: Georgia Tech Traces OLMo Abilities Back to Training Data(3 posts)→
More from Research
- Retriever: A Framework for Asynchronous, Closed-Loop Robot Agents — ZeYanjie · 2026-08-24
- Converting GMMs ↔ PEFs for fast KLD approximation — FrnkNlsn · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24