TIDE distills diffusion data attribution into millisecond lookups, 4-5 orders of magnitude cheaper
serrjoa · x · 2026-10-09
A Sony team (Shixuan Liu et al.) posted arXiv:2609.38776 on efficient training data attribution for diffusion models. They formulate attribution via a local score discrepancy measure applicable to any diffusion variant (DDPM, EDM, flow matching), estimable as a preconditioned gradient similarity without retraining. TID uses Kronecker-factored curvature to avoid random projections and per-sample gradient storage; it's distilled into TIDE, a forward-only student reproducing the teacher's rankings from internal activations. On CIFAR-10, ArtBench-10, and MS-COCO counterfactual evals, TID matches or beats SOTA while TIDE retains most accuracy at 4-5 orders of magnitude lower per-query cost—attributing generated samples in milliseconds, faster than generation itself. Code coming soon.
More from Research
- Building Rome from a single image: new method reconstructs 3D scenes with geometry beyond the visible view — jonstephens85 · 2026-10-09
- AI plus formal tools compress years of research into hours as community live-tracks integer multiplication breakthrough — aran_nayebi · 2026-10-09
- PredActor unifies humanoid motion generation and control in one diffusion policy at 50Hz on Unitree G1 — carlosdponx · 2026-10-09
- Off-policy training works for continual learning if done right, new thread argues — AdtRaghunathan · 2026-10-09
- sudo L7 benchmark: best coding agents pass only ~45% of staff-level tasks — echen · 2026-10-09
- Google's FlowAgent fixes CI test failures: 67% correct on 195 cases, 28,554 fixes applied after launch — omarsar0 · 2026-10-09