
Gregory Kurtzer has launched OpenWALDO, an open-source AI training dataset containing 167.3 billion tokens from public sources like government records and academic papers.
The project seeks to address the opacity of most open-weight models, which use closed training data of unknown origin, by creating a shared, auditable foundation that companies can build upon with their own proprietary data.
While the dataset is far smaller than the trillions of tokens used by frontier AI models, Kurtzer argues this transparency model mirrors the success of open-source software and could improve efficiency across the industry.
What happened
Gregory Kurtzer, founder of CentOS and Rocky Linux, has launched OpenWALDO (Open Weights, Artifacts, Licenses, Data, Origins), an open-source AI training dataset funded by his company CIQ. The project currently contains 167.3 billion reference tokens drawn from government records, academic papers, mailing lists, and public domain literature.
Why it matters
Most open-weight AI models use closed-source training data of unknown origin and licensing, creating risks for customers and wasting computing resources through duplicate industry efforts. OpenWALDO aims to make training data transparent and auditable, letting companies add their own proprietary data on top of a verified public baseline—a model Kurtzer argues mirrors how open-source software succeeded by being inspectable and community-validated rather than proprietary.
What to watch
The 167.3 billion tokens in OpenWALDO represent a small fraction compared to the tens of trillions used to train frontier models. Adoption by AI companies will determine whether the project becomes a shared industry foundation or remains a niche effort; CIQ did not confirm whether any model has yet been trained on the dataset.
Gregory Kurtzer, the founder of CentOS and Rocky Linux, has announced OpenWALDO, a new open-source AI training dataset initiative funded by his company CIQ, which also sponsors Rocky Linux. The project name stands for Open Weights, Artifacts, Licenses, Data, Origins, and reflects Kurtzer's effort to apply open-source principles to AI model development—an area he argues has yet to embrace true openness despite the recent prominence of open-weight models.
The core problem OpenWALDO addresses is that most downloadable open-weight models, while their code is publicly available, rely on closed-source training data. This data often combines copyrighted material, distilled responses from other proprietary models, and user-generated content obtained without clear consent or license disclosure. CIQ argues this opacity creates multiple risks: customers have no way to audit what trained their model or under what licenses, hidden data sources can taint model behavior and put software stacks at risk, and the industry wastes computing resources by each organization separately gathering and processing training data. A single shared public baseline would eliminate this duplication and allow improvements to benefit all future models trained on it.
Kurtzer positions OpenWALDO as a solution that mirrors the success of open-source software. Companies and labs can take the OpenWALDO corpus as a verified baseline, add their own proprietary data, build their model, and maintain a clear auditable trail back to the source of every component. This approach addresses concerns about security and misuse that frontier labs have raised about open-weight models; Kurtzer counters that open-source software faced identical skepticism but won through inspectability, forkability, and community validation—not safety claims. "Linux didn't win by being certified safe. It won by being inspectable, forkable, and community validated," he said. "AI is missing that same property, and OpenWALDO is how we build it."
The timing reflects industry momentum behind open-weight models, particularly from China, which are closing in on the capabilities of closed-source frontier models like ChatGPT and Claude. Many businesses now question why they should pay for AI services they don't own, cannot truly control, and have no visibility into. The OpenWALDO dataset currently contains 167.3 billion reference tokens gathered from government records, open-source academic papers, mailing lists, and public domain literature—a substantial amount, but a fraction of the tens of trillions used to train frontier models. CIQ did not confirm whether any organization has yet trained a model on the dataset. Adoption remains uncertain; while the project reflects legitimate industry pain points, persuading competitors to share training work represents a major cultural shift for an industry accustomed to treating data as a proprietary advantage.
OpenWALDO addresses a growing tension in the AI industry: despite the rise of downloadable open-weight models, their training data remains opaque. Companies and labs typically draw from copyrighted material, distilled responses from other models, or user-generated content obtained without clear consent—sources that are neither disclosed nor auditable. This opacity creates two problems the project aims to solve. First, it poses practical risk: customers cannot know whether a model's foundation is tainted, putting their software stacks in jeopardy. Second, it wastes resources: without a shared baseline, every organization duplicates the effort of curating and validating training data.
Kurtzer's approach borrows from open-source software's playbook. Rather than keeping the dataset proprietary or claiming safety through obscurity, OpenWALDO makes sources traceable and forkable. Companies can then layer their own proprietary data on top of the public baseline, creating a clear audit trail from final model back to its foundations. This model has already resonated with businesses questioning why they pay for closed AI services they cannot control or inspect, especially as open-weight models from China begin matching the capabilities of proprietary alternatives like ChatGPT and Claude.
However, OpenWALDO faces a steep climb. The 167.3 billion tokens currently available are a drop next to the tens of trillions used by frontier labs, and it remains unclear whether any model has been trained on the dataset yet. Adoption requires the AI industry to abandon its closed-data practices—a shift that took decades for open-source software and may prove harder in an AI landscape where companies view proprietary training data as a core competitive advantage.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Eli Lilly and Company has committed to roll out Veeva Vault CRM across its global operations, adopting Veeva's…

Visa, the global payments network, is positioning itself to profit from AI-driven commerce by securing its rol…

Oracle Health released a redesigned patient portal that includes new AI-powered capabilities

McDonald's CEO Chris Kempczinski announced on August 4 that the company is consolidating customer data from 70…

S&P Global announced an expanded partnership with Microsoft to embed its proprietary data and analytics direct…

S&P Global announced an expanded collaboration with Microsoft to integrate its data, insights and analytics di…

The AI news that matters, in one minute each morning.
Sign up free