Paper Overview
Research Area: Computer Vision Authors: Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Kai Huang, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu Published: 2026-09-18 arXiv: 2609.22069
Abstract
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce.
To address these gaps, the authors introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models.
Key Contributions
- OmniVBench benchmark: expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions — covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings.
- Element-anchored evaluation: includes 12,172 use-case-specific checks evaluating whether expected reference elements are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to instructions.
- Omni-R2V Dataset: sourced mainly from a large-scale professional video footage library, containing 340,000 processed training samples, providing an industrial-scale, diverse training resource for the research community.
- Data construction pipeline: task-specific reference-target pair construction offering a practical, scalable path to building omni R2V data.
Findings
Large-scale evaluation of state-of-the-art open-source and closed-source R2V models reveals clear performance gaps across OmniVBench task families and evaluation dimensions, highlighting the limitations of current R2V models.
---
*Auto-collected on 2026-09-22*