Friday, 04 September 2026 PDT | 04:07 PM
The 1 News Alt Logo Text Smart News for Global Indians

‘Attribution decay’ complicates the picture of AI

AI News September 05, 2026 04:00 AM
‘Attribution decay’ complicates the picture of AI

Art generated with artificial intelligence (AI) relies on models trained on billions of images. That training has been the subject of much debate and several ongoing or impending lawsuits. Many consider the training to be a form of involuntary extraction and that the images derived from it are a violation of countless artists’ copyrights, but recent scientific findings about how AI processes the information it is trained on complicate this view.

Scientists at the Massachusetts Institute of Technology’s Computer Science and Artificial Intelligence Laboratory published new research in Nature Communications on 18 August that explores a phenomenon they have termed “attribution decay”. The researchers, Zheng Dai and David K. Gifford, found that in large datasets, removing certain data did not change the output. The argument that could be made in light of this goes something like: if a copyrighted image was part of the training data, and removing it did not change the generated image, then the generated image could not be accused of copyright infringement.

The researchers “developed a method for taking away one piece of the training sets and then regenerating the image as though that piece of training data didn’t exist”, Dai says. “And if you find that [the output] doesn’t change much, then you can’t attribute it to that piece of data, because it didn’t have any influence on the final output.” He adds that in the context of image generation, "when you train on large data sets, there is no piece of data that you can take out that significantly alters the image, and that [leads] to this idea of unattributability”.

So for example, if A Bigger Splash (1967) and even David Hockney’s entire oeuvre were removed from a dataset, could the AI tool still churn out a Bigger Splash-like image with the right set of prompts? Dai says that such scenarios would be speculative, but if the training data contained traces of A Bigger Splash, the AI tool could still generate something resembling the famous painting at Tate Britain. In other words, even if all direct references to A Bigger Splash are removed, the dataset might still include enough indirect references, derivative images like parodies, glimpses of it or echoes of it to generate a Bigger Splash-like image. In this hypothetical scenario, Hockney's original painting would not be part of the dataset the generation relied on, so one could argue in court that there was no copyright infringement.

“If you view [the] training off of people’s data as [infringement],” Dai says, “it is sort of ironic, the more infringement you do, the less infringing the output is.”

Framing their findings as scientific evidence that theft is OK, so long as it is done at a massive scale, is unlikely to win anyone over to the pro-AI side of the argument. Although AI tools can allow for direct infringement when prompted to replicate specific source material, the public sentiment against such models does not account for the sheer magnitude of the datasets involved. When these algorithms generate images based on broader prompts, they are referencing datasets comprising millions if not billions of images. To use a crude metaphor, if every online image of an artwork created by an artist were a grain of sand, to accuse an AI artist of violating your copyright each time they generate an image is a bit like walking up to a stranger on a beach and accusing them of building a sandcastle with your sand.

A loophole in this clumsy metaphor is that the systems at work are capable of choosing precisely which grains of sand to use if prompted to do so. More importantly, there is a big difference between holding an AI user accountable for the outputs their prompts generate, and holding the AI company accountable for how they trained their model.

Asked if his and Gifford’s findings effectively serve as a get-out-of-jail-for-free card for tech companies being accused of copyright infringement over their AI models, Dai has a caveat. “It certainly looks like it from one angle, but I think it is trickier than that,” he says. “For example, if every single artist joined a class action lawsuit and said, ‘Oh no, you can’t use [our works]’... If you take out every single piece of artwork [from a dataset], then obviously the model wouldn't exist.” But, he added that in a legal battle, companies could remove a copyrighted image to show its absence does not change the output, demonstrating unattributability as an argument against copyright infringement.

It seems unlikely that a global artists’ movement will lobby to collectively opt out of all AI training, but even if we did, AI models would still be able to draw from works that are in the public domain or derivative works, and a plethora of additional types of material we may not have even considered. Dai and Gifford’s findings do not exonerate the training process itself; even if an AI company were to argue for unattributability in a particular output, it could still be held liable for how its algorithms were trained.

Attribution decay complicates an already contentious legal landscape in which the public perception is often that all generated output violates copyright, and none of it is copyrightable. To be sure, AI-generated works of art can still be copyrighted under certain conditions. The United States Copyright Office insists on a case-by-case approach that looks for the existence of human creative expression.