LAION drops massive open video dataset with 10 million hours of footage for AI research
LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research, containing 10 million hours of footage scraped from CommonCrawl The dataset comprises 80 million videos, 55 million clips with auto-generated video and audio descriptions, and 300 million still images, primarily sourced from YouTube and mostly in English Models trained on BVD outperform comparable InternVid-trained models by up to 2.1 percentage points on standard video-to-text benchmarks The data
Analysis
TL;DR
- LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research, containing 10 million hours of footage scraped from CommonCrawl
- The dataset comprises 80 million videos, 55 million clips with auto-generated video and audio descriptions, and 300 million still images, primarily sourced from YouTube and mostly in English
- Models trained on BVD outperform comparable InternVid-trained models by up to 2.1 percentage points on standard video-to-text benchmarks
- The dataset is released for research-only use, backed by a 2024 Hamburg Regional Court ruling permitting non-commercial collection of copyrighted content
- Both the dataset and accompanying code are freely available for the AI research community
Why It Matters
The release of BVD addresses a critical bottleneck in multimodal AI development: the scarcity of large-scale, open video datasets compared to text and image resources. For researchers building video-language models, this dataset provides an unprecedented scale of synchronized visual, audio, and textual data that can significantly advance capabilities in video understanding, generation, and cross-modal reasoning.
Technical Details
- Scale and Composition: 10 million hours of video from 80 million downloaded URLs out of 1.3 billion found in CommonCrawl, yielding 55 million clips with auto-generated descriptions and 300 million extracted still images
- Multimodal Training Approach: The dataset supports joint training across video, audio, and text modalities, enabling models to learn correspondences between visual content, audio signals, and natural language descriptions
- Benchmark Performance: BVD-trained models achieve up to 2.1 percentage point improvements over InternVid-trained counterparts on common video-to-text evaluation benchmarks
- Legal Framework: Released under research-only restrictions, with legal grounding from a 2024 Hamburg Regional Court decision allowing non-commercial collection of copyrighted material
Industry Insight
- The availability of a dataset at this scale for video AI research could accelerate the development of more capable multimodal foundation models, potentially narrowing the gap between video and text/image model capabilities
- The research-only licensing model sets an important precedent for how large-scale web-scraped datasets can be legally distributed, likely influencing future dataset releases in the multimodal space
- Researchers and organizations should prioritize integrating BVD into their training pipelines to remain competitive on video understanding benchmarks, while carefully adhering to the non-commercial usage restrictions
Disclaimer: The above content is generated by AI and is for reference only.