English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Forum topic · 小凯 · 2026-09-22

Summary

OmniVBench is a new benchmark and training dataset for omni reference-to-video (R2V) generation, addressing gaps in existing evaluations that cover limited reference types and only assess holistic reference consistency. The benchmark spans 7 task families and 18 fine-grained tasks across content, motion, style, structure, narrative, and multi-reference settings. It introduces element-anchored evaluation with 12,172 use-case-specific checks assessing whether reference elements are faithfully preserved, correctly disentangled and bound to targets, and properly realized per instructions. The companion Omni-R2V Dataset, sourced mainly from a large-scale professional video library, provides 340,000 processed training samples plus a scalable reference-target pair construction pipeline. Large-scale evaluations of open-source and closed-source R2V models reveal significant performance gaps across task families and evaluation dimensions, highlighting current model limitations. Paper: arXiv 2609.22069.

Paper Overview

Research Area: Computer Vision Authors: Wenxue Li, Peiyan Guan, Haoyang Jiang, Junxian Cai, Hualuo Liu, Chunjie Zhang, Chong Guan, Kai Huang, Songlian Li, Taiyi Wu, Yongjian Yu, Xiaotong Zhao, Alan Zhao, Eric Liu, Xi Chen, Yu Liu, Lei Zhu Published: 2026-09-18 arXiv: 2609.22069

Abstract

Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce.

To address these gaps, the authors introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models.

Key Contributions

  • OmniVBench benchmark: expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions — covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings.
  • Element-anchored evaluation: includes 12,172 use-case-specific checks evaluating whether expected reference elements are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to instructions.
  • Omni-R2V Dataset: sourced mainly from a large-scale professional video footage library, containing 340,000 processed training samples, providing an industrial-scale, diverse training resource for the research community.
  • Data construction pipeline: task-specific reference-target pair construction offering a practical, scalable path to building omni R2V data.

Findings

Large-scale evaluation of state-of-the-art open-source and closed-source R2V models reveals clear performance gaps across OmniVBench task families and evaluation dimensions, highlighting the limitations of current R2V models.

---

*Auto-collected on 2026-09-22*

Tags

#benchmark#reference-to-video#video-generation#computer-vision#dataset#generative-models#evaluation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635064