• Research from MIT CSAIL shows that when generative AI models are trained on massive datasets, tracing an output back to a specific original work is nearly impossible.
  • The research team calls this phenomenon “attribution decay,” where the ability to attribute an output to specific training data decreases as the data scale increases.
  • Experiments show that one image, an artist’s entire portfolio, or all photos of a person can be removed from the training data with almost no change to the output.
  • According to the team, if removing a piece of data leaves the result unchanged, that data cannot be considered the cause of the output.
  • They developed a “diffusion ensemble” architecture, dividing the model into small components to precisely remove the influence of individual data without full retraining.
  • The team compared the new system with 24 traditional diffusion models on the same data, achieving equivalent image quality.
  • The trial used over 160,000 images from 7 public datasets, including CIFAR-10, CelebA, MetFaces, and ArtBench.
  • Results show that the influence of individual data points decreases according to an inverse power law as data scale grows.
  • The team trained 1,282 separate models for verification, obtaining consistent results.
  • Professor David Gifford suggests this finding reopens the question of whether AI output remains a derivative work or qualifies for copyright protection.
  • The research also suggests AI companies could build models ensuring outputs cannot be attributed to specific works.
  • The authors note the study currently applies only to image-generating diffusion models, not yet confirmed for large language models.

📌 MIT’s work presents a new perspective on generative AI copyright. When training data is large enough, individual works have almost no significant impact on the generated image, making it very difficult to identify the “source author.” This finding could heavily impact AI copyright lawsuits, how derivative works are assessed, and compensation regulations for creators, though it remains unproven for large language models.

Share.
VIET NAM CONSULTING AND MEASUREMENT JOINT STOCK COMPANY
Contact

Email: info@vietmetric.vn
Address: No. 34, Alley 91, Tran Duy Hung Street, Yen Hoa Ward, Hanoi City

© 2026 Vietmetric
Exit mobile version