[论文] GraphVid: Interactive Graph-Controllable Video Generation

研究领域: CV 作者: Vedant Shah, Onkar Susladkar, Tushar Prakash 发布时间: 2025-07-27 arXiv: 2507.21741

论文概要

研究领域: CV 作者: Vedant Shah, Onkar Susladkar, Tushar Prakash 发布时间: 2025-07-27 arXiv: 2507.21741

中文摘要

可控视频生成仍然具有挑战性,因为难以使用文本提示或主要约束像素运动的运动控制输入来指定精确的多物体交互。在实践中,基于轨迹的控制通常要求用户为多个物体绘制精确轨迹,这随着场景复杂度增加而难以扩展,且在遮挡或重叠下变得模糊。为实现灵活而精确的多主体控制,我们引入了GraphVid,一种图条件图像到视频生成模型,通过结构化交互图实现交互控制。我们进一步策划了GraphVid-Bench,一个大规模以交互为中心的视频数据集,具有结构化关系注释,用于训练交互感知的视频生成模型。尽管使用的训练数据和可训练参数远少于先前的运动控制方法,GraphVid仍提供了强大的可控性和视频质量。与Motion-I2V相比,GraphVid将FID降低多达39.9%,FVD降低37.6%,同时提高了PSNR(9.87=>15.98)和SSIM(0.38=>0.61)。我们的结果突显了结构化语义界面作为可控视频生成的强大范式的潜力。

原文摘要

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene复杂度 and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using subst...


*自动采集于 2026-07-27*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens