• Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms & Conditions
Monday, September 21, 2026
freevirtualsolutions.com
  • Home
  • Virtual Assistant
  • Enterprise Technology
  • Customer Support
  • Marketing Support
  • Workforce Management
No Result
View All Result
  • Home
  • Virtual Assistant
  • Enterprise Technology
  • Customer Support
  • Marketing Support
  • Workforce Management
No Result
View All Result
Free Virtual Solutions
No Result
View All Result
Home Enterprise Technology

AI models get convenient amnesia about source material as they grow, MIT boffins find

Ermias S. by Ermias S.
1 month ago
in Enterprise Technology
0
399
SHARES
2.3k
VIEWS
Share on FacebookShare on Twitter



ai and ml

Attributing diffusion model output to a specific input becomes more difficult with more training data

The process of training an AI model becomes a paradox at scale – the more it remembers, the less it remembers about the source of its memories.

MIT computer scientists went looking for a way to attribute AI model output to specific training data, in the hope that understanding could inform AI regulation. 

What they found, described in a paper titled, “Outputs of Generative Diffusion Models are Often Unattributable,” looks like it will actually make regulation more difficult. Scientific journal Nature Communications will publish the paper on Tuesday.

The authors, Zheng Dai and David K Gifford, affiliated with MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL), note that diffusion models like Midjourney and Stable Diffusion have become widely used tools for generating artifacts including images, videos, and audio.

Diffusion models have also attracted lawsuits from artists who argue that copies of their work included in training data have enabled AI models to reproduce their output and artistic style. In one such ongoing copyright case from 2023, Andersen et al. v. Stability AI Ltd, the plaintiffs have been trying to convince the court to make defendant Midjourney provide the datasets used to train its models.

The plaintiffs allege that Midjourney made its training datasets “by scraping images associated with specific artists’ names for the express purpose of enabling its model to mimic those artists’ expressive content.”

Being able to attribute model output to the content they ingested during training would help people understand how models function and would have various applications “including machine unlearning, data poisoning, model interpretability, fairness, and privacy,” Dai and Gifford wrote in their paper. 

“Furthermore, given the contemporary adoption of these models for creative and commercial purposes, attributability also carries ethical, policy, financial, and legal implications.”

But as it turns out, attributing model output to a specific input becomes more difficult as models get larger.

“Here we show that attribution, characterized as the task of locating a part of the training data that can be held responsible for a generated sample, can become impossible if a model is trained on a sufficiently large corpus of data,” the authors state.

“We find that the more data a model is trained on, the less attributable its generated samples become, a phenomenon we henceforth refer to as attribution decay.”

The authors tested this by removing specific training data through a process called ablation. And the result of this testing showed that for very large models, you could take away, for example, the image of the Mona Lisa or all of Leonardo Da Vinci’s work — yet the model could still reproduce that image or style.

In MIT’s press release, Gifford argues that the findings suggest models are creative in the sense that they’re not just copying their training data.

“If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn’t attributable to anything on the internet,” he said.

He also argues that having the ability to test whether a given output is derivative creates an obligation for companies to show their work cannot be attributed to a particular source. 

At the same time, the research suggests a liability avoidance strategy – make your model large enough that no output can be attributed to any one specific input.

James Grimmelmann, a law professor at Cornell Law School and Cornell Tech, said in a statement, “If attribution worked, it would reliably tell us whether similarities between a model’s output and a copyright-protected work are due to copying or coincidence. But this paper provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying.”

In an email to The Register, Grimmelmann said current copyright claims against AI companies haven’t focused specifically on the extent to which images similar to artist’s real-world work can be elicited from a model.

“German cases have dealt with apparently memorized outputs from music models, and US cases have dealt with whether training is a fair use, but output similarity for image models remains mostly untested in court,” he said. ®



Source_link

Tags: amnesiaboffinsconvenientFindGrowmaterialMITmodelsSource
Previous Post

How Demand Gen Teams Spot and Close Competitive Gaps Before Deals Slip Away

Next Post

China-Linked Hackers Use AI Agents to Launch Autonomous Cyberattack on Taiwan

Next Post
AI Crawlers Are Reshaping the Web—And DataDome Is Drawing the Line

China-Linked Hackers Use AI Agents to Launch Autonomous Cyberattack on Taiwan

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Categories

  • Customer Support (1,762)
  • Enterprise Technology (4,563)
  • Marketing Support (3,885)
  • Virtual Assistant (4,147)
  • Workforce Management (1,535)

Free Virtual Solutions

Welcome to freevirtualsolutionsThe goal of freevirtualsolutions is to give you the absolute best news sources for any topic! Our topics are carefully curated and constantly updated as we know the web moves fast so we try to as well.

Category

  • Customer Support
  • Enterprise Technology
  • Marketing Support
  • Virtual Assistant
  • Workforce Management

Recent Post

  • Top Advantages of Moving Beyond a PEO in 2026
  • Schneider says hotter coolant can make AI datacenters less thirsty
  • How WFM Software Can Positively Impact a Contact Center Culture
  • Home
  • About Us
  • Contact Us
  • Disclaimer
  • Privacy Policy
  • Terms & Conditions

Copyright © 2023 Freevirtualsolutions.com | All Rights Reserved.

No Result
View All Result
  • Home
  • Virtual Assistant
  • Enterprise Technology
  • Customer Support
  • Marketing Support
  • Workforce Management

Copyright © 2023 Freevirtualsolutions.com | All Rights Reserved.

Chat with us

Hi there! How can I help you?