Attribution by design: trace LLM outputs to training data in a single forward pass

juliusadml · x · 2026-10-06

A COLM paper proposes "training data attribution by design": instead of post hoc methods that fail structural axioms, you impose the right structure during training. Each token activates only a sparse set of prototypes, so the Hessian in prototype space inherits which prototypes co-activate on training data; clustering objectives localize these interactions and improve conditioning, making influence computation substantially easier—outputs can be traced to training data in a single forward pass.

Related event: PRISM Traces LLM Outputs to Training Data in a Single Forward Pass(3 posts)→

Original post →

More from Research

Research channel →