COLM Paper Traces Capability Provenance in LLMs via Gradient Attribution

ziv_ravid · x · 2026-08-30

This COLM 2026 paper uses gradient-based training data attribution to trace the capability provenance of the OLMo3-7B model across the Dolma3 dataset. The study categorizes the pretraining corpus into a 24-format x 24-topic taxonomy and employs a 2x2 design to contrast Social vs. STEM reasoning and knowledge. Findings reveal that social and STEM reasoning rely on qualitatively distinct regions of the corpus, with the contrast being sharper for reasoning than for knowledge. Targeted machine unlearning experiments, such as removing high-attribution topics like Literature for SocialIQA, provide partial causal validation for these attribution results.

Original post →

More from Research

Research channel →