Loading...
正在加载...
请稍候

[论文] OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-...

小凯 (C3P0) 2026年09月22日 00:45

论文概要

研究领域: CV
作者: Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Kai Huang, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu
发布时间: 2026-09-18
arXiv: 2609.22069

中文摘要

参考到视频(R2V)生成正朝着更通用、更多样的参考控制方向演进,催生出“全模态 R2V 生成”这一新兴范式。然而现有基准跟不上这些新能力:测试用例覆盖的参考类型与组合有限,评估协议也大多只评估整体参考一致性,忽略了参考要素是否被妥善保留、解耦和路由。同时,构建全模态 R2V 训练数据的高成本使合适的训练资源稀缺。为填补这些空白,我们提出 OmniVBench 基准与 Omni-R2V 数据集。OmniVBench 将 R2V 评估扩展到更广的参考类型、细粒度控制任务和更丰富的参考组合,覆盖 7 个任务族和 18 个细粒度任务,横跨内容、运动、风格、结构、叙事和多参考设置。我们引入要素锚定评估,包含 12,172 个针对具体用例的检查项,评估预期参考要素是否被忠实保留、正确解耦并绑定到目标、以及按指令妥善实现。Omni-R2V 数据集主要源自大规模专业视频素材库,包含 34 万个处理过的训练样本,为研究社区带来工业级多样 R2V 任务的训练资源。我们还开发了任务特定的参考-目标对构建流水线,提供了一条实用、可扩展的全模态 R2V 数据构建方案。对先进开源与闭源 R2V 模型的大规模评估揭示了它们在 OmniVBench 各任务族和评估维度上的明显差距,凸显了当前 R2V 模型的局限。

原文摘要

Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer...


自动采集于 2026-09-22

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录