Most AI-generated images trained on large datasets can’t be traced back to the data they were trained on, potentially throwing a curveball in intellectual property theft cases, a new study from MIT’s Computer Science and Artificial Intelligence Laboratory found.
Researchers discovered a phenomenon they call “attribution decay,” where the more data a generative model is trained on, the harder it becomes to trace a generated image to a single image from the training data — all Picasso’s work could be removed from the training data, for example, and the AI-generated image might still resemble a Picasso.
AI companies have used books, articles, photos, and other works, often without consulting authors, to train their models, sparking a flurry of copyright infringement lawsuits. Disney, NBCUniversal, and DreamWorks filed an IP lawsuit last year against AI image-generator MidJourney; The New York Times sued OpenAI and Microsoft in 2023 for a similar reason. The study raises legal questions about training data and fair use policy, according to the authors.
“We might have to rethink what intellectual property means,” Zheng Dai, a former MIT researcher and lead author of the work, told Semafor. “You can’t just assume it, and the attribution link sort of vanishes.”



