Attribution by design: trace LLM outputs to training data in a single forward pass
juliusadml · x · 2026-10-06
A COLM paper proposes "training data attribution by design": instead of post hoc methods that fail structural axioms, you impose the right structure during training. Each token activates only a sparse set of prototypes, so the Hessian in prototype space inherits which prototypes co-activate on training data; clustering objectives localize these interactions and improve conditioning, making influence computation substantially easier—outputs can be traced to training data in a single forward pass.
Related event: PRISM Traces LLM Outputs to Training Data in a Single Forward Pass(3 posts)→
More from Research
- Force Prompting: video generation models learn physics-based control signals — creatoroff · 2026-10-07
- MIT CSAIL's DREAM unifies image understanding and generation, ~10% faster — MIT_CSAIL · 2026-10-07
- E2S finetuning turns off-policy expert data into on-policy student samples via amortized sampling — Lianhuiq · 2026-10-06
- IdeaLens: a new detector that tries to tell whether a text's ideas came from a human or AI — MohitIyyer · 2026-10-06
- PRISM: trace an LLM's outputs back to training data in a single forward pass — juliusadml · 2026-10-06
- COLM 2026 paper acceptance statistics released — xwang_lk · 2026-10-06