English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VEFX-Bench: A Holistic Benchmark and Reward Model for Instruction-Guided Video Editing

Forum topic · 小凯 · 2026-04-21

Summary

VEFX-Bench is a benchmark suite from researchers at Texas A&M University (arXiv 2604.16272) addressing the lack of large-scale, human-annotated data and standardized evaluation for instruction-guided video editing. The team introduces VEFX-Dataset, 5,049 human-annotated video editing examples spanning 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: instruction following, rendering quality, and edit exclusivity. Building on this, they propose VEFX-Reward, a reward model that jointly processes the source video, the editing instruction, and the edited video, predicting per-dimension quality scores via ordinal regression. They also release VEFX-Bench, 300 curated video-prompt pairs for standardized comparison of editing systems. Experiments show VEFX-Reward aligns better with human judgments than generic vision-language model judges and prior reward models on IQA/VQA metrics and group preference evaluations. Benchmarking commercial and open-source video editing systems reveals persistent gaps in visual plausibility, instruction following, and edit locality.

Overview

As AI-assisted video creation becomes practical, instruction-guided video editing is increasingly essential for refining generated or captured footage to meet professional requirements. However, the field has lacked both a large-scale human-annotated dataset with complete editing examples and a standardized evaluator for comparing editing systems. Existing resources suffer from small scale, missing edited outputs, or absent human quality labels, while evaluation typically relies on costly manual inspection or generic VLM judges not specialized for editing quality.

  • Paper: arXiv 2604.16272
  • Authors: Xiangbo Gao, Sicong Jiang, Bangya Liu, Xinghao Chen, Minglai Yang, Siyuan Yang, Mingyang Wu, Jiongze Yu, Qi Zheng, Haozhi Wang, Jiayi Zhang, Jared Yang, Jie Yang, Zihan Wang, Qing Yin, Zhengzhong Tu
  • Key Contributions

    1. VEFX-Dataset: A human-annotated dataset of 5,049 video editing examples across 9 major editing categories and 32 subcategories. Each example is labeled along three decoupled dimensions:

  • Instruction following
  • Rendering quality
  • Edit exclusivity (whether only the intended content changed)
  • 2. VEFX-Reward: A reward model specialized for video editing quality assessment. It jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores using ordinal regression.

    3. VEFX-Bench: A benchmark of 300 curated video-prompt pairs enabling standardized comparison of editing systems.

    Findings

  • VEFX-Reward aligns better with human judgment than generic VLM judges and previous reward models, on both standard IQA/VQA metrics and group preference evaluations.
  • Using VEFX-Reward as an evaluator, benchmarking of representative commercial and open-source video editing systems reveals persistent gaps in visual plausibility, instruction following, and edit locality.

Why It Matters

The combination of a large human-labeled dataset, a dedicated reward model, and a standardized benchmark provides the community with the tools needed to measure and improve instruction-guided video editing systems systematically rather than relying on expensive manual review or ill-suited general-purpose judges.

Tags

#video-editing#benchmark#reward-model#ai-generated-content#evaluation#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618609