AIToday

Tech leaders back open AI, but deeper openness needed—data access key

Hacker News4h agoSend on LINE
Tech leaders back open AI, but deeper openness needed—data access key

Key takeaway

Jensen Huang and nearly all major U.S. tech firms have signed a letter backing open-source AI development, citing national interest and innovation benefits. However, the author argues this does not go far enough: for open models to truly compete with proprietary ones, the U.S. should also guarantee open-source developers fair-use access to the same copyrighted material—books, papers, code, videos—that closed models were trained on, a shift that would require rethinking intellectual property law.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Jensen Huang and CEOs across the tech industry (all firms except Anthropic) signed a letter Friday supporting U.S. investment in open-source artificial intelligence, responding to Trump administration signals about restricting open-source software as a protectionist move.

  • Why it matters

    Open-source AI models could level the playing field in economic competition with China and accelerate innovation through shared research and benchmarking—but the letter does not address the hardest part: giving open models fair-use access to the same training data (books, papers, code, videos) that proprietary models use, which would require rethinking copyright and intellectual property law.

  • What to watch

    The author argues that true competitive open-source models require three things—people, compute, and data—and that Chinese companies like DeepSeek and Moonshot have shown models can be trained affordably (reportedly under ten million dollars), suggesting compute agreements among tech firms are feasible if data access is resolved.

In Depth

On Friday, Jensen Huang published a short letter on Twitter calling for American investment in "a strong, open ecosystem" around open-source artificial intelligence. Huang argued that open source had "created a shared foundation of knowledge on which generations of American engineers and entrepreneurs built their institutional sovereignty," and predicted similar benefits would flow from embracing open-source AI. The letter carried implicit political weight: the Trump administration had been signaling it would ban open-source software as a protectionist move to suppress Chinese language models, a policy that would have benefited Anthropic and OpenAI by eliminating open competitors. Recognizing this, CEOs from across the tech industry—all major and smaller firms except Anthropic—quickly signed on to Huang's letter on Friday.

The author of this post celebrates the development but argues Huang left out the most difficult and important piece. Building a competitive language model requires three inputs: people, compute, and data. On people, the author notes that although the promise of wealth has lured talent away from open-source work, industry backing should now stabilize recruitment. On compute, the author challenges the hyperscaler narrative that massive datacenters are an insurmountable moat: Chinese companies DeepSeek and Moonshot trained competitive models for reportedly under ten million dollars, and open-source development could be supported through compute-sharing agreements among tech firms, which would represent only a minuscule fraction of each firm's budget. The author predicts that once models are built openly, efficiency will follow naturally: "machine learning is a field that innovates through frictionless reproducibility," and openly available research, code, and data—evaluated through competitive benchmarking—rapidly improve systems. The author cites the example of high-quality ImageNet models, which went from trainable only at Google to buildable on a desktop in less than a year.

The hard part, the author argues, is data. Proprietary AI companies have trained large language models on vast quantities of copyrighted material: pirated libraries of books and academic papers, copyrighted imagery, collaborative knowledge bases like Wikipedia, volunteer forums like Reddit, all public code on GitHub, and transcripts of every video posted to YouTube. These companies have been found liable in court and have admitted this practice in their published papers, claiming it falls under fair use because training is transformative. The author accepts this legal reasoning but argues it cuts both ways: if closed-model training on copyrighted material is fair use, then open-source distillation of proprietary model outputs should also be fair use, and open-source models should have equal fair-use access to the same training corpora. This "open corpus" would require a fundamental rethinking of intellectual property law and copyright, but the author frames it as essential for competitive parity. The author concludes that while the challenges are large—recruiting talent, securing compute collaboration, and overhauling copyright—they are surmountable, and that the tech industry's support, combined with the brief history of open-source dominance before the proprietary 2020s, makes this the moment to "open things up again."

Context & Analysis

Jensen Huang's Friday letter represents a major shift: the Trump administration's protectionist stance against open-source software unexpectedly unified nearly the entire tech industry (with Anthropic as the sole holdout) in support of open AI—a reversal that signals both geopolitical anxiety about Chinese AI competitiveness and recognition that openness itself drives innovation. The author endorses this momentum but argues it is incomplete. The letter addresses open weights and open source, but sidesteps the data problem: proprietary AI labs have trained models on vast quantities of copyrighted material—books, academic papers, code repositories, videos—often without explicit permission, citing transformative fair use as legal cover. Open-source developers, by contrast, lack comparable access to these datasets, creating an asymmetry that prevents fair competition. The author frames this as a moral and strategic question: if closed models can legally use pirated and copyrighted material for training, why should open models not have the same fair-use right, especially in service of national technological sovereignty against China? The author acknowledges three hard challenges—recruiting talent away from private AI companies, securing compute infrastructure through industrial collaboration (citing Chinese examples like DeepSeek and Moonshot that trained models for under ten million dollars), and overhauling copyright law—but argues all are surmountable given industry consensus and the precedent of open-source innovation in earlier decades.

FAQ

What did Jensen Huang's letter say?
Huang called on Twitter for American investment in "a strong, open ecosystem" around open-source artificial intelligence, arguing that open source "created a shared foundation of knowledge on which generations of American engineers and entrepreneurs built their institutional sovereignty," and predicting the same would happen with open-source AI.
Why did all these tech CEOs sign so quickly?
The Trump administration had been signaling it would ban open-source software as a protectionist move to suppress Chinese language models; CEOs signed on because they saw that such a ban would benefit only Anthropic and OpenAI.
What does the author say is still missing?
The author argues open-source models need fair-use access to the same training data (pirated books, academic papers, copyrighted images, Wikipedia, Reddit, GitHub code, YouTube transcripts) that proprietary models used, calling this the "open corpus"—and that this would require rethinking intellectual property and copyright law.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime