[论文] LIFT: Layout-In-Future Video Generation under Large Viewpoint Change v...
研究领域: CV 作者: Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuo…
论文概要
研究领域: CV 作者: Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu 发布时间: 2026-09-29 arXiv: 2609.38146
中文摘要
我们提出LIFT——一个统一的图像到视频生成框架,在相机控制之外补充了'未来布局'(Layout-In-FuTure)控制,使用户能够指定未来视图中应出现什么内容以及出现在哪里。这解决了可控视频生成中的一个实际需求:给定初始图像,用户不仅关心相机如何移动,还关心场景在关键未来时刻(尤其是最终帧)应呈现什么样子。现有的相机控制只能指定视角轨迹,而文本提示仅提供粗略的语义引导;两者都无法精确决定未来视图的内容和空间布局。这一限制在大视角变化下尤为突出——此时相机会揭示第一帧中不可见的区域。因此,LIFT使用末帧布局作为期望未来场景的显式控制信号。由于从如此稀疏的布局引导中学习远比在密集的每帧布局条件上学习更具挑战性,我们引入在线策略自蒸馏(OPSD)将密集布局教师的控制能力迁移到末帧布局学生。我们还构建了LIFT-Vista数据集,其中包含大视角变化的相机和时间一致的布局标注。实验表明,LIFT在视频质量、未来布局可控性和相机可控性上均优于其他方法。
原文摘要
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first fr...
*自动采集于 2026-10-01*
#论文 #arXiv #CV #小凯