Technology newsletter icon
From Semafor Technology
In your inbox, 2x per week
Sign up

Gen AI outputs are unattributable, study finds

Aug 19, 2026, 3:02pm EDT
PostEmailWhatsapp
Top left: an image from a model trained on public-domain artwork by 744 artists. Beside it: the alternate versions produced if each artist, in turn, had been excluded from training. Dai, Z., Gifford, D.K. Outputs of generative diffusion models are often unattributable.

Most AI-generated images trained on large datasets can’t be traced back to the data they were trained on, potentially throwing a curveball in intellectual property theft cases, a new study from MIT’s Computer Science and Artificial Intelligence Laboratory found.

Researchers discovered a phenomenon they call “attribution decay,” where the more data a generative model is trained on, the harder it becomes to trace a generated image to a single image from the training data — all Picasso’s work could be removed from the training data, for example, and the AI-generated image might still resemble a Picasso.

AI companies have used books, articles, photos, and other works, often without consulting authors, to train their models, sparking a flurry of copyright infringement lawsuits. Disney, NBCUniversal, and DreamWorks filed an IP lawsuit last year against AI image-generator MidJourney; The New York Times sued OpenAI and Microsoft in 2023 for a similar reason. The study raises legal questions about training data and fair use policy, according to the authors.

“We might have to rethink what intellectual property means,” Zheng Dai, a former MIT researcher and lead author of the work, told Semafor. “You can’t just assume it, and the attribution link sort of vanishes.”

AD