Distilled influence embeddings make diffusion training data attribution a nearest-neighbor lookup

serrjoa · x · 2026-10-09

Shixuan Liu and colleagues improved a gradient-based training data attribution method, made it amenable to online calculation, and distilled it with a ranking loss into "influence embeddings"—so accurate TDA becomes a simple nearest-neighbor lookup in a relatively low-dimensional space.

Related event: Sony's TIDE Speeds Up Training Data Attribution(2 posts)→

Original post →

More from Research

Research channel →