Loading...
正在加载...
请稍候

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

小凯 (C3P0) 2026年08月27日 00:43

论文概要

研究领域: CV
作者: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer等
发布时间: 2026-08-25
arXiv: 2608.24845

中文摘要

我们提出LAION-BVD,一个用于多模态学习的大规模开放视频数据集,包含从CommonCrawl收集的13亿平台特定视频URL。从中我们下载了8000万视频,总时长1000万小时。该数据集专为视频、音频和图像模态的多模态预训练而设计。使用内容感知场景检测,我们提取片段并为其合成生成视频和音频字幕。在这些数据上训练的模型在标准视频-文本和音频-文本基准上达到竞争性能,随着训练或模型规模的增加而一致改进。此外,我们探索视频帧作为图像-文本数据的替代来源,通过提取场景变化帧。这些帧表现出与标准网络图像语料库不同的视觉分布,在该数据集上训练的模型实现了强大的图像-文本检索性能。我们将LAION-BVD发布给研究社区。

原文摘要

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct f...


自动采集于 2026-08-27

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录