MIT study finds "attribution decay": removing any single artist's data often changes generated images not at all

MIT_CSAIL · x · 2026-08-19

New research from MIT CSAIL identifies a phenomenon called attribution decay: the more data a generative model is trained on, the less any individual training example matters to any particular output. At sufficient scale, the researchers find you can often remove any single image, every image by a given artist, or every photograph of a given person—and the generated sample doesn't change.

Their argument: if removing a piece of data leaves the output unchanged, that data didn't affect the output, so attributing the output to it makes little sense—and doing this for every piece of data one at a time suggests no single item is responsible. The team also found AI-generated images hard to trace to specific training data. The work raises fundamental questions for ongoing copyright lawsuits, licensing deals, and proposed regulations: for large-scale models, the question of "whose work went into this image" may often have no answer.

Related event: MIT Finds "Attribution Decay" in AI-Generated Art(2 posts)→

Original post →

More from Safety

Safety channel →