
Gregory Kurtzer, the founder of Rocky Linux, has launched OpenWALDO, an open-source initiative to create a shared, auditable corpus of AI training data that is transparent about sources and licenses.
The project addresses a major gap in the AI industry, where the MIT-led Data Provenance Initiative found that over 70% of datasets omit license information and over 50% contain license errors.
OpenWALDO's public corpus currently contains 124 billion reference tokens across 75.1 million documents and aims to make training data as inspectable and verifiable as open-source software, allowing developers to build on a common base layer rather than repeatedly assembling basic training data from scratch.
What happened
Gregory Kurtzer, founder of Rocky Linux and co-founder of CentOS, released OpenWALDO, an open-source project designed to make AI training data transparent and auditable. The project's public corpus currently indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers.
Why it matters
AI training data has remained largely opaque and proprietary, with the MIT-led Data Provenance Initiative finding license-omission rates above 70% and license-error rates greater than 50% across 1,800+ text datasets it audited. OpenWALDO proposes making training data behave like open-source dependencies—named, reviewable, versioned, and attributable—so developers can verify sources and avoid problems like recursive training on outputs from other AI systems (model collapse).
What to watch
OpenWALDO's success depends on sustained supply of high-quality, legally usable data; trusted review and governance processes; and whether AI developers will adopt its provenance records as a standard part of model-development workflows. Kurtzer's company CIQ will support the project commercially rather than control it.
Gregory Kurtzer, the founder of Rocky Linux and co-founder of CentOS, has launched OpenWALDO, a community-governed open-source initiative aimed at bringing transparency and auditability to AI training data. The project's full name—Open Weights, Artifacts, Licenses, Data and Origins—reflects its mission to make training data behave like open-source software dependencies: named, reviewable, versioned, attributable, and verifiable. Backed by Kurtzer's enterprise infrastructure company CIQ, OpenWALDO addresses what Kurtzer calls a critical gap in the AI industry, where training data has remained largely closed and unexamined even as open-weights models and open-source AI code have become more prevalent.
The technical approach centers on three mechanisms: Git-based review for metadata, content-addressed storage for underlying data objects, and an AI Bill of Materials (ABOM) that tracks resolved data and model lineage through training and release. Rather than building its own foundation model, OpenWALDO provides a common base layer—a licensed, provenance-tracked training corpus that model builders can use, extend with proprietary data, and carry into their own training runs. This structure is designed to let organizations focus resources on differentiation rather than repeatedly assembling basic training datasets. At launch, OpenWALDO's public corpus indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers.
The timing reflects growing concern about training data provenance and licensing. The MIT-led Data Provenance Initiative audited more than 1,800 text datasets and found that license information was often omitted or miscategorized, with license-omission rates above 70% and error rates greater than 50%. OpenWALDO also addresses a quality problem: recursive training on outputs from other AI systems, known as model collapse, degrades model performance. By maintaining a public record of training sources, developers can distinguish and select training material more carefully. The project adopts practices familiar from software development—public review, attributable contributions, Developer Certificate of Origin sign-offs, and preserved provenance from data ingestion through release—and will accept contributors from individuals, researchers, institutions, and companies, while CIQ provides commercial support rather than control.
Kurtzer framed the initiative in terms of open-source philosophy: "I've spent my career watching open source turn users into builders, competitors into collaborators, and shared problems into common infrastructure that operates at massive scale." He argued that "Linux didn't win by being certified safe. It won by being inspectable, forkable, and community validated. AI is missing that same property, and OpenWALDO is how we build it." Success will depend on OpenWALDO's ability to attract sustained supplies of high-quality, legally usable data; establish trusted review and governance processes; and persuade AI developers that its provenance records provide enough practical value to become standard in model-development workflows.
The AI industry has long struggled with opacity around training data sources and licenses. The MIT-led Data Provenance Initiative's audit of over 1,800 text datasets revealed a systemic problem: license information was omitted in over 70% of cases and miscategorized in more than 50%. This matters because models trained on data with unclear or incorrect licensing create legal and ethical risks for organizations that deploy them, and because training on outputs from other AI systems without proper attribution can lead to model collapse—a degradation in model quality caused by recursive training on synthetic data.
Kurtzer's proposal frames OpenWALDO as the missing layer in the "open AI" movement. While open-weights models (those whose parameters are publicly available) and open-source AI code have become more common, the training data itself has remained largely proprietary and unauditable. OpenWALDO introduces an AI Bill of Materials (ABOM) that tracks data lineage, licensing, and sources through training and release—practices borrowed from software development. The project positions this as shared infrastructure that would allow organizations to spend resources on differentiation rather than repeatedly sourcing and validating basic training datasets from scratch. Kurtzer explicitly draws a parallel to Linux's success, arguing that inspectability, forkability, and community validation—not certification—drove adoption.
The real test will be adoption. OpenWALDO must attract sustained contributions of high-quality, legally usable data; establish governance processes that command trust; and demonstrate that provenance records provide practical value in model-development workflows. CIQ will provide commercial support but not control the project, a structure intended to preserve its community character.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
NVIDIA and partners released multiple open-source AI models optimized for local execution throughout August, i…

Leica Biosystems announced multiple FDA 510(k) clearances on August 11, 2026, including Aperio iQC DX software…

Anthropic is embedding imperceptible, machine-readable watermarks into text generated by Claude models release…

A directory of 27 creators—writers, editors, software engineers, musicians, and designers—has been published…
Security researchers led by Alexander Panfilov discovered a vulnerability in the APIs of all major AI provider…

Apple is developing an iOS feature called Apple Reference Image that embeds provenance metadata into iPhone ph…

The AI news that matters, in one minute each morning.
Sign up free