Influence Matching improves dataset distillation on Tiny-ImageNet and Flickr30K
TheHKU · hf · 2026-07-24
The paper Dataset Distillation by Influence Matching revisits dataset distillation from an outcome-centric view.
- Instead of matching training trajectories or per-step gradients, it matches the final effect on converged parameters.
- The method introduces a differentiable, sample-level influence estimator that measures how adding or removing data shifts parameters.
- It avoids inverse-Hessian products and convexity assumptions by unrolling optimization dynamics with a first-order Taylor approximation.
- The synthetic set is learned by minimizing the mismatch between the influence of real data and the synthetic set.
Results reported in the abstract include:
- Tiny-ImageNet (IPC=10): 31.5%, beating NCFM by +4.7%
- Vision-language distillation on Flickr30K, with 200–1000 synthetic samples, outperforming strong process-matching baselines by +2.5% on average image/text retrieval results.
Code is planned for release at the linked GitHub repo.
More from Research
- Large university study finds ChatGPT had no detectable effect on grades after COVID disruption was modeled — emollick · 2026-07-25
- Physics-IQ audit finds ambiguous prompts and artifacts can reshuffle video-model rankings — DynamicWebPaige · 2026-07-25
- Podcast preview: HUG explores human universal grasping for robots — chris_j_paxton · 2026-07-25
- New preprint uses AlphaFold3 ensembles to predict TCR–pMHC binding — quaidmorris · 2026-07-25
- Stanford HAI warns that averaging expert safety scores can erase good chatbot advice — StanfordHAI · 2026-07-25
- Claude Opus 4.8 hits 16.5% on ProgramBench’s almost-resolved metric — jyangballin · 2026-07-25