[论文] ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
研究领域: CV 作者: Xinghao Chen, Xiangbo Gao, Jiongze Yu, et al. 发布时间: 2026-09-30 arXiv: 2609.26755
论文概要
研究领域: CV 作者: Xinghao Chen, Xiangbo Gao, Jiongze Yu, et al. 发布时间: 2026-09-30 arXiv: 2609.26755
中文摘要
近期视频生成越来越逼真和可控,但视频编辑仍然发展不足,尤其是对于必须保留原始场景动态的精确局部编辑。视频场景文本编辑替换场景表面上的文本——如店面招牌、白板和产品标签——同时保留周围内容、运动和相机动态。虽然图像场景文本编辑已被充分研究,但同时实现高视觉质量、时间一致性和编辑局部性的视频场景文本编辑仍未被充分探索。现有资源提供的有配对真实视频数据有限,且通用的视频编辑指标不能直接衡量被请求的文本是否随时间保持正确。我们引入 ViTeX-Bench,一个包含 ViTeX 数据集和三轴评估协议的基准套件。数据集包含 387 个真实世界 720p 视频,带有文本区域掩码和编辑指令:其中 230 个提供经审核的流水线生成配对编辑用于训练,157 个构成冻结评估集。该协议通过 13 个指标评估文本正确性、视觉和时间质量以及编辑局部性,每个轴有一个主要指标,并用帕累托比较分析其权衡。OCR 校准、人工评估和标注敏感性分析支持对这些分数的解释。在来自四个编辑家族的八个基线方法中,准确的文本、时间稳定性和场景保留仍然难以同时实现。我们还发布了 ViTeX-Edit-14B,一个在配对训练集上微调的开源参考编辑器,采用运动对齐的字形-视频条件化。它实现了 0.688 的字符准确率,是评估的视频原生编辑器中最高的平均值,并且在原始编辑器输出中取得了最低的可比文本裁剪 Warp。
原文摘要
Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing is well studied for images, video scene text editing that achieves high visual quality, temporal consistency, and edit locality remains underexplored. Existing resources offer limited paired real-video data, and general video-editing metrics do not directly measure whether the requested text remains correct over time. We introduce ViTeX-Bench, a benchmark suite comprising ViTeX-Dataset ...
*自动采集于 2026-10-02*
#论文 #arXiv #CV #小凯