AI News AI资讯 5h ago Updated 1h ago 更新于 1小时前 50

LAION drops massive open video dataset with 10 million hours of footage for AI research LAION发布大规模开源视频数据集,包含1000万小时影像供AI研究

LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research, containing 10 million hours of footage scraped from CommonCrawl The dataset comprises 80 million videos, 55 million clips with auto-generated video and audio descriptions, and 300 million still images, primarily sourced from YouTube and mostly in English Models trained on BVD outperform comparable InternVid-trained models by up to 2.1 percentage points on standard video-to-text benchmarks The data LAION发布Big Video Dataset (BVD),包含1000万小时视频,是目前最大开源视频数据集之一 从CommonCrawl的13亿视频URL中下载8000万视频,提取5500万片段(含自动生成视频/音频描述)及3亿静态图像 BVD训练的模型在视频到文本基准测试上比InternVid训练的模型高出2.1个百分点 数据集仅用于研究目的,基于2024年汉堡地方法院裁决收集版权内容,数据集和代码免费开放 数据主要来自YouTube,以英语内容为主

72
Hot 热度
68
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • LAION released the Big Video Dataset (BVD), one of the largest open video datasets for AI research, containing 10 million hours of footage scraped from CommonCrawl
  • The dataset comprises 80 million videos, 55 million clips with auto-generated video and audio descriptions, and 300 million still images, primarily sourced from YouTube and mostly in English
  • Models trained on BVD outperform comparable InternVid-trained models by up to 2.1 percentage points on standard video-to-text benchmarks
  • The dataset is released for research-only use, backed by a 2024 Hamburg Regional Court ruling permitting non-commercial collection of copyrighted content
  • Both the dataset and accompanying code are freely available for the AI research community

Why It Matters

The release of BVD addresses a critical bottleneck in multimodal AI development: the scarcity of large-scale, open video datasets compared to text and image resources. For researchers building video-language models, this dataset provides an unprecedented scale of synchronized visual, audio, and textual data that can significantly advance capabilities in video understanding, generation, and cross-modal reasoning.

Technical Details

  • Scale and Composition: 10 million hours of video from 80 million downloaded URLs out of 1.3 billion found in CommonCrawl, yielding 55 million clips with auto-generated descriptions and 300 million extracted still images
  • Multimodal Training Approach: The dataset supports joint training across video, audio, and text modalities, enabling models to learn correspondences between visual content, audio signals, and natural language descriptions
  • Benchmark Performance: BVD-trained models achieve up to 2.1 percentage point improvements over InternVid-trained counterparts on common video-to-text evaluation benchmarks
  • Legal Framework: Released under research-only restrictions, with legal grounding from a 2024 Hamburg Regional Court decision allowing non-commercial collection of copyrighted material

Industry Insight

  • The availability of a dataset at this scale for video AI research could accelerate the development of more capable multimodal foundation models, potentially narrowing the gap between video and text/image model capabilities
  • The research-only licensing model sets an important precedent for how large-scale web-scraped datasets can be legally distributed, likely influencing future dataset releases in the multimodal space
  • Researchers and organizations should prioritize integrating BVD into their training pipelines to remain competitive on video understanding benchmarks, while carefully adhering to the non-commercial usage restrictions

TL;DR

  • LAION发布Big Video Dataset (BVD),包含1000万小时视频,是目前最大开源视频数据集之一
  • 从CommonCrawl的13亿视频URL中下载8000万视频,提取5500万片段(含自动生成视频/音频描述)及3亿静态图像
  • BVD训练的模型在视频到文本基准测试上比InternVid训练的模型高出2.1个百分点
  • 数据集仅用于研究目的,基于2024年汉堡地方法院裁决收集版权内容,数据集和代码免费开放
  • 数据主要来自YouTube,以英语内容为主

为什么值得看

BVD为多模态AI研究提供了大规模开源视频数据,填补了视频理解领域的关键空白。其性能提升证明了大规模视频-音频-文本联合训练的价值,对视频生成、理解和多模态模型发展具有重要意义。

技术解析

  • 数据来源与规模:从CommonCrawl的13亿视频URL中筛选,下载8000万视频,总时长达1000万小时;提取5500万视频片段,每个片段附带自动生成的视频和音频描述,同时生成3亿静态图像
  • 性能表现:在常见视频到文本基准测试中,BVD训练的模型比InternVid训练的模型高出2.1个百分点
  • 训练方式:将视频、音频和文本三种模态联合训练,学习视觉内容与文本描述及声音的对应关系
  • 法律框架:基于2024年汉堡地方法院裁决,允许为非商业研究目的收集版权内容;LAION强调用户需尊重原创作者权益
  • 开放许可:数据集和代码完全免费开放,仅限研究用途

行业启示

  • 开源视频数据集的突破将显著降低多模态AI研究门槛,加速视频理解、视频生成等方向的研究进展
  • 大规模视频-音频-文本联合训练的有效性得到验证,为未来多模态模型架构设计提供了重要参考方向
  • 法律框架的明确为AI研究使用版权内容提供了先例,但行业仍需建立更完善的版权合规机制以平衡创新与创作者权益

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Open Source 开源 Dataset 数据集 Video Generation 视频生成 Multimodal 多模态 Research 科学研究