论文概要
研究领域: CV
作者: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer等
发布时间: 2026-08-25
arXiv: 2608.24845
中文摘要
我们提出LAION-BVD,一个用于多模态学习的大规模开放视频数据集,包含从CommonCrawl收集的13亿平台特定视频URL。从中我们下载了8000万视频,总时长1000万小时。该数据集专为视频、音频和图像模态的多模态预训练而设计。使用内容感知场景检测,我们提取片段并为其合成生成视频和音频字幕。在这些数据上训练的模型在标准视频-文本和音频-文本基准上达到竞争性能,随着训练或模型规模的增加而一致改进。此外,我们探索视频帧作为图像-文本数据的替代来源,通过提取场景变化帧。这些帧表现出与标准网络图像语料库不同的视觉分布,在该数据集上训练的模型实现了强大的图像-文本检索性能。我们将LAION-BVD发布给研究社区。
原文摘要
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct f...
自动采集于 2026-08-27
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。