
LAION released one of the largest open video datasets. It contains 10 million hours and 55 million clips.
Models trained on it beat InternVid-trained ones by up to 2.1 points.
The dataset is for research only.
What happened
LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research. It includes 10 million hours of video, 55 million clips with auto-generated descriptions, and 300 million still images.
Why it matters
Models trained on BVD outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. The dataset is available for research only, and LAION's legal basis may rest on a 2024 Hamburg court ruling allowing copyrighted content for non-commercial research.
What to watch
The dataset and code are freely available, but users are asked to respect creators' rights. The data draws from 80 million downloaded videos out of 1.3 billion URLs found in CommonCrawl, mostly from YouTube and predominantly in English.
Ask the AI about this article →
LAION's release of the Big Video Dataset is notable not just for its size but for the training results it enables. According to the accompanying paper, models trained on this data perform better than those trained on InternVid by up to 2.1 percentage points on standard video-to-text benchmarks. This suggests that the scale and diversity of the dataset—drawing from 80 million videos out of 1.3 billion URLs—can improve AI's ability to connect visuals with descriptions and audio.
The dataset's research-only restriction reflects a careful legal stance. LAION appears to rely on a 2024 Hamburg court ruling that permitted collecting copyrighted material for non-commercial research. While this may shield the organization, it also means commercial applications would need to seek separate licensing or use other data sources. The request to respect creators' rights highlights the ongoing tension between AI data needs and copyright.
For business readers, the relevance lies in what this enables: better video understanding models that could eventually power products like search, translation, or automated captioning. However, the research-only limitation means any commercial use must wait for a clearer legal path or proprietary alternatives. The dataset's predominantly English content also means gains may be uneven across languages.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Numerous large language models (LLMs, AI that understands and generates text) from big names like Meta and Goo…

Meta Platforms released Muse Glimmer, a 30-billion-parameter AI model with open weights, on Aug

Nvidia is reported to be acquiring Hugging Face for $13 billion, following its $6 billion deal with Poolside a…

Anthropic PBC previewed the Model Hardware Standard (MHS), a unified interface that lets AI agents control sci…
NVIDIA has agreed to acquire AI platform Hugging Face for $2 trillion, according to the article

Nvidia has reportedly agreed to acquire Hugging Face, the AI project hosting platform, for $12.9 billion