AIToday
Image GenerationAI Safety & AlignmentAI Regulation & PolicySemafor TechPublished: Aug 20, 2026, 06:01 JST3 min read

MIT study: AI images untraceable to training sources

MIT study: AI images untraceable to training sources

Key takeaway

  • A new MIT study has found that large generative AI models cannot be reliably traced back to specific images in their training data — a phenomenon called "attribution decay" where larger datasets make it harder to establish which original works influenced a generated image.

  • The discovery poses a direct challenge to ongoing copyright lawsuits against companies like OpenAI, Microsoft, and MidJourney, which trained models on copyrighted works without permission, because it may become nearly impossible to prove that a generated image infringes a particular artist's copyright if the attribution link cannot be established.

3 Key Points

  1. What happened

    Researchers at MIT's Computer Science and Artificial Intelligence Laboratory found that generative AI models trained on large datasets cannot reliably be traced back to individual images in their training data — a phenomenon they call "attribution decay." The more data a model is trained on, the harder it becomes to link a generated image to a specific source, even if a particular artist's entire body of work were removed from the training set.

  2. Why it matters

    This finding complicates copyright enforcement. AI companies including MidJourney, OpenAI, and Microsoft have trained models on books, articles, photos, and other works often without author permission, triggering lawsuits from Disney, NBCUniversal, DreamWorks, and The New York Times. The study's lead author, Zheng Dai (a former MIT researcher), notes that the inability to establish attribution links may force a rethinking of intellectual property law and fair use policy — "the attribution link sort of vanishes" — making it harder to prove infringement in court.

  3. What to watch

    The legal implications. Copyright cases now underway will likely hinge on whether courts accept that attribution decay is a fundamental property of large-scale generative models, potentially reshaping how intellectual property claims are evaluated against AI training practices.

Ask the AI about this article →

Context & Analysis

The study addresses a critical gap exposed by the recent wave of copyright litigation against AI companies. Since 2023, major studios and publishers have sued AI developers for training models on copyrighted material without authorization. These cases rest on the ability to demonstrate that a generated output is derived from a specific protected work — but the MIT researchers have found that this causal link becomes effectively invisible as models scale. The phenomenon of attribution decay suggests that training on larger datasets actually obscures the provenance of generated outputs, making it harder for plaintiffs to establish the precise copyright violation needed to win in court.

This discovery reframes the intellectual property debate. Previous lawsuits assumed that tracing generated content back to its training sources would be a technical and legal tool to prove infringement. The study indicates that assumption may not hold. Lead researcher Zheng Dai explicitly raises the prospect that courts and regulators may need to fundamentally redefine intellectual property rights in the age of large-scale AI training — moving away from a model based on demonstrable attribution of individual works and toward a different framework entirely.

FAQ

What is attribution decay?
Attribution decay is a phenomenon where the more data a generative AI model is trained on, the harder it becomes to trace a generated image back to a specific image in the training data. For example, even if all of Picasso's works were removed from the training data, an AI-generated image might still resemble a Picasso.
Who filed copyright lawsuits against AI companies?
Disney, NBCUniversal, and DreamWorks filed an IP lawsuit last year against MidJourney. The New York Times sued OpenAI and Microsoft in 2023 for using their works to train AI models without permission.

Get the latest Image Generation news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleModel Downloads Become a Geopolitical Weapon—and Your Supply Chain Risk

The AI news that matters, in one minute each morning.

Sign up free