论文概要
研究领域: CV
作者: Byungjun Kim, Taeksoo Kim, Hyunsoo Cha, Hanbyul Joo
发布时间: 2026-07-24
arXiv: 2607.22535
中文摘要
动作条件化视频世界模型根据初始观测和动作信号预测未来观测。在机器人学中,动作通过两个不同过程影响未来观测:首先由机器人体和控制器将动作实现为机器人运动,然后通过接触和物体运动使场景响应。直接以动作命令为条件要求世界模型学习实现过程本身,而以记录的未来状态为条件则泄露了它本应预测的交互结果。本文提出机器人分解世界模型,将两个机器人特定因素移出世界模型。首先是动作实现:每个命令通过机器人自身的控制器和运动学滚动为可部署的名义轨迹,这是一种中间信号,既避免了动作实现学习,也避免了未来状态泄露。其次是机器人渲染:该名义轨迹通过机器人URDF进行渲染,将机器人的几何、运动学和外观从模型中分解出来,转化为显式渲染的机器人几何体。为解决深度歧义,我们将末端执行器深度与场景深度配对,提供超越图像平面重叠的几何线索来判断接触和遮挡。相机感知的静态RGB/深度上下文与渲染的机器人几何体共同构成一个共享的视觉世界模型接口,该接口在不同视角和机器人体现之间保持一致,因此模型将动作仅视为可见的机器人几何体,并学习物体如何对其响应。实验表明,渲染接口优于向量条件化基线,并在推理时泛化到未见的机器人体现。本文进一步展示了该模型通过将手部运动重定向并渲染为机器人几何体,从人类演示生成机器人操作视频。
原文摘要
Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence future observations through two distinct processes: they are first realized into robot motion by the robot body and controller, and the scene then responds through contact and object motion. Conditioning directly on action commands asks the world model to learn the realization process itself, while conditioning on logged future states leaks the interaction outcomes it is meant to predict. We propose robot-factored world models, which move two robot-specific factors outside the world model. First, action realization: each command is rolled through the robot's own controller and kinematics into a deployment-available nominal trajectory, a middle s...
自动采集于 2026-07-28
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。