AIToday

Microsoft trained MAI models on unlicensed web data, contradicting clean-data pledge

THE DECODERJun 5, 2026
Microsoft trained MAI models on unlicensed web data, contradicting clean-data pledge

Key takeaway

Microsoft's technical paper shows its MAI models were trained partly on unlicensed web data from sources like Common Crawl, contradicting the company's earlier promise that training used only "enterprise grade, clean and commercially licensed data." The company relies on fair use and a proprietary crawler that respects robots.txt—an approach shared by other AI firms—but marketed as uniquely clean. This gap between stated practice and actual training sources matters to businesses evaluating Microsoft's AI credibility against competitors.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Microsoft's technical paper reveals that its new MAI models were partly trained on unlicensed web data from sources including Common Crawl, despite the company's earlier claim that they used only "enterprise grade, clean and commercially licensed data." Microsoft says it uses a proprietary crawler that respects the Robots Exclusion Protocol (robots.txt) and related controls to manage how content is accessed.

  • Why it matters

    Microsoft marketed its training data as especially clean, but the paper shows it relies on the open internet like other AI companies do. The company relies on a fair-use argument—a legal framework still being contested and sorted out in courts—to justify web scraping. This suggests businesses considering Microsoft's AI offerings should not assume the underlying training data is fundamentally different from competitors'.

  • What to watch

    The paper describes the data as a "mixture of publicly available and licensed human-generated data," placing the burden of protecting content on individual site owners through robots.txt and HTML controls rather than explicit licensing agreements.

FAQ

What sources did Microsoft use to train its MAI models?
Microsoft used a mixture of publicly available and licensed human-generated data, including Common Crawl. The company also deployed a proprietary crawler that respects the Robots Exclusion Protocol (robots.txt) and related meta-tag and HTML controls to manage how content on sites is accessed.
What did Microsoft originally claim about its training data?
Microsoft had previously claimed the MAI models were trained only on "enterprise grade, clean and commercially licensed data."

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →