AIToday
AI Safety & AlignmentOpen-Source AIHacker NewsPublished: Aug 12, 2026, 06:00 JST5 min read

Rocky Linux founder launches OpenWALDO to bring transparency to AI training data

Rocky Linux founder launches OpenWALDO to bring transparency to AI training data

Key takeaway

  • Gregory Kurtzer, the founder of Rocky Linux, has launched OpenWALDO, an open-source initiative to create a shared, auditable corpus of AI training data that is transparent about sources and licenses.

  • The project addresses a major gap in the AI industry, where the MIT-led Data Provenance Initiative found that over 70% of datasets omit license information and over 50% contain license errors.

  • OpenWALDO's public corpus currently contains 124 billion reference tokens across 75.1 million documents and aims to make training data as inspectable and verifiable as open-source software, allowing developers to build on a common base layer rather than repeatedly assembling basic training data from scratch.

3 Key Points

  1. What happened

    Gregory Kurtzer, founder of Rocky Linux and co-founder of CentOS, released OpenWALDO, an open-source project designed to make AI training data transparent and auditable. The project's public corpus currently indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers.

  2. Why it matters

    AI training data has remained largely opaque and proprietary, with the MIT-led Data Provenance Initiative finding license-omission rates above 70% and license-error rates greater than 50% across 1,800+ text datasets it audited. OpenWALDO proposes making training data behave like open-source dependencies—named, reviewable, versioned, and attributable—so developers can verify sources and avoid problems like recursive training on outputs from other AI systems (model collapse).

  3. What to watch

    OpenWALDO's success depends on sustained supply of high-quality, legally usable data; trusted review and governance processes; and whether AI developers will adopt its provenance records as a standard part of model-development workflows. Kurtzer's company CIQ will support the project commercially rather than control it.

In Depth

Read the full story

Gregory Kurtzer, the founder of Rocky Linux and co-founder of CentOS, has launched OpenWALDO, a community-governed open-source initiative aimed at bringing transparency and auditability to AI training data. The project's full name—Open Weights, Artifacts, Licenses, Data and Origins—reflects its mission to make training data behave like open-source software dependencies: named, reviewable, versioned, attributable, and verifiable. Backed by Kurtzer's enterprise infrastructure company CIQ, OpenWALDO addresses what Kurtzer calls a critical gap in the AI industry, where training data has remained largely closed and unexamined even as open-weights models and open-source AI code have become more prevalent.

The technical approach centers on three mechanisms: Git-based review for metadata, content-addressed storage for underlying data objects, and an AI Bill of Materials (ABOM) that tracks resolved data and model lineage through training and release. Rather than building its own foundation model, OpenWALDO provides a common base layer—a licensed, provenance-tracked training corpus that model builders can use, extend with proprietary data, and carry into their own training runs. This structure is designed to let organizations focus resources on differentiation rather than repeatedly assembling basic training datasets. At launch, OpenWALDO's public corpus indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers.

The timing reflects growing concern about training data provenance and licensing. The MIT-led Data Provenance Initiative audited more than 1,800 text datasets and found that license information was often omitted or miscategorized, with license-omission rates above 70% and error rates greater than 50%. OpenWALDO also addresses a quality problem: recursive training on outputs from other AI systems, known as model collapse, degrades model performance. By maintaining a public record of training sources, developers can distinguish and select training material more carefully. The project adopts practices familiar from software development—public review, attributable contributions, Developer Certificate of Origin sign-offs, and preserved provenance from data ingestion through release—and will accept contributors from individuals, researchers, institutions, and companies, while CIQ provides commercial support rather than control.

Kurtzer framed the initiative in terms of open-source philosophy: "I've spent my career watching open source turn users into builders, competitors into collaborators, and shared problems into common infrastructure that operates at massive scale." He argued that "Linux didn't win by being certified safe. It won by being inspectable, forkable, and community validated. AI is missing that same property, and OpenWALDO is how we build it." Success will depend on OpenWALDO's ability to attract sustained supplies of high-quality, legally usable data; establish trusted review and governance processes; and persuade AI developers that its provenance records provide enough practical value to become standard in model-development workflows.

Context & Analysis

The AI industry has long struggled with opacity around training data sources and licenses. The MIT-led Data Provenance Initiative's audit of over 1,800 text datasets revealed a systemic problem: license information was omitted in over 70% of cases and miscategorized in more than 50%. This matters because models trained on data with unclear or incorrect licensing create legal and ethical risks for organizations that deploy them, and because training on outputs from other AI systems without proper attribution can lead to model collapse—a degradation in model quality caused by recursive training on synthetic data.

Kurtzer's proposal frames OpenWALDO as the missing layer in the "open AI" movement. While open-weights models (those whose parameters are publicly available) and open-source AI code have become more common, the training data itself has remained largely proprietary and unauditable. OpenWALDO introduces an AI Bill of Materials (ABOM) that tracks data lineage, licensing, and sources through training and release—practices borrowed from software development. The project positions this as shared infrastructure that would allow organizations to spend resources on differentiation rather than repeatedly sourcing and validating basic training datasets from scratch. Kurtzer explicitly draws a parallel to Linux's success, arguing that inspectability, forkability, and community validation—not certification—drove adoption.

The real test will be adoption. OpenWALDO must attract sustained contributions of high-quality, legally usable data; establish governance processes that command trust; and demonstrate that provenance records provide practical value in model-development workflows. CIQ will provide commercial support but not control the project, a structure intended to preserve its community character.

FAQ

What exactly is OpenWALDO, and what does the name mean?
OpenWALDO stands for Open Weights, Artifacts, Licenses, Data and Origins. It is a community-governed open-source project designed to create a shared, auditable corpus of AI training data and code that tracks provenance and licensing information.
How much training data does OpenWALDO currently have?
OpenWALDO's public corpus currently indexes 124 billion reference tokens across 75.1 million documents, organized into 20 corpora and 1,051 shards, with 18 asserted license identifiers.
Will OpenWALDO build its own AI model?
No. OpenWALDO does not plan to build a foundation model. Instead, it aims to provide a common base layer—a licensed, provenance-tracked training corpus that model builders can use, extend with proprietary material, and carry into their own training runs.

Get the latest AI Safety & Alignment news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleResearcher rebuilds spiking language model around CPU inference

The AI news that matters, in one minute each morning.

Sign up free