AIToday
Open-Source AITHE DECODERPublished: Aug 29, 2026, 19:02 JST2 min read

LAION releases 10M-hour open video dataset

LAION releases 10M-hour open video dataset

Key takeaway

  • LAION released one of the largest open video datasets. It contains 10 million hours and 55 million clips.

  • Models trained on it beat InternVid-trained ones by up to 2.1 points.

  • The dataset is for research only.

3 Key Points

  1. What happened

    LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research. It includes 10 million hours of video, 55 million clips with auto-generated descriptions, and 300 million still images.

  2. Why it matters

    Models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. The dataset is available for research only, and LAION's legal basis may rest on a 2024 Hamburg court ruling allowing copyrighted content for non-commercial research.

  3. What to watch

    The dataset and code are freely available, but users are asked to respect creators' rights. The data draws from 80 million downloaded videos out of 1.3 billion URLs found in CommonCrawl, mostly from YouTube and predominantly in English.

Ask the AI about this article →

Context & Analysis

LAION's release of the Big Video Dataset is notable not just for its size but for the training results it enables. According to the accompanying paper, models trained on this data perform better than those trained on InternVid by up to 2.1 percentage points on standard video-to-text benchmarks. This suggests that the scale and diversity of the dataset—drawing from 80 million videos out of 1.3 billion URLs—can improve AI's ability to connect visuals with descriptions and audio.

The dataset's research-only restriction reflects a careful legal stance. LAION appears to rely on a 2024 Hamburg court ruling that permitted collecting copyrighted material for non-commercial research. While this may shield the organization, it also means commercial applications would need to seek separate licensing or use other data sources. The request to respect creators' rights highlights the ongoing tension between AI data needs and copyright.

For business readers, the relevance lies in what this enables: better video understanding models that could eventually power products like search, translation, or automated captioning. However, the research-only limitation means any commercial use must wait for a clearer legal path or proprietary alternatives. The dataset's predominantly English content also means gains may be uneven across languages.

FAQ

How large is the dataset?
It includes 10 million hours of video, 55 million clips with auto-generated descriptions, and 300 million still images, sourced from 80 million downloaded videos out of 1.3 billion URLs found in CommonCrawl.
Who can use this dataset?
It is released for research only. LAION asks users to respect the rights of original content creators, and it is freely available along with the code.
Why does LAION think it can use copyrighted content?
The organization can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research.

Get the latest Open-Source AI news every morning

For example, today's edition would include:

  • Chatbot on Your Own PC: Privacy, Offline, FreeWIRED AI · 1h ago
  • Meta open-sources Muse Glimmer, keeps Muse Spark 1.2 closedYahoo Finance AI · 4h ago
  • Nvidia's reported $13B Hugging Face deal highlights open-weight AI M&A surgeTechCrunch AI · 19h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleReddit Stock Down 36.7%, Nears Fair Value