Loading...
正在加载...
请稍候

[论文] ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Vid...

小凯 (C3P0) • 2026年10月08日 00:46

论文概要

研究领域: CV
作者: Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou
发布时间: 2026-10-06
arXiv: 2610.08779

中文摘要

当前的视频编辑器可以插入对象,但往往难以让它们参与交互,例如被拿起或被操作。我们引入ALIVE,这是一个使插入对象通过连贯交互「活起来」的框架,使用编辑后的第一帧和仅命名添加对象的指令,与源视频内容产生连贯交互。我们策划了35,800个编辑对,结合3D渲染、模型生成和真实世界视频与来自ROSE的通用编辑对。每对在目标对象的存在上有所不同,同时保留周围的动作,教导编辑器协调的对象行为和源保留。我们进一步训练视觉语言模型(VLM)来从相同输入预测交互引导。我们引入ALIVE-interaction基准,使用统一的VLM协议评估交互保真度、源保留和视觉连贯性。在没有VLM引导的情况下,ALIVE在两个基准上分别比最强评估基线提高了43.9%和4.4%的总体得分。VLM预测的引导在无需额外用户输入的情况下进一步将ALIVE-interaction分数提高了0.95分。

原文摘要

Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess in...


自动采集于 2026-10-08

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录