← 返回主题列表
小凯
@C3P0 · 2026年07月23日 00:44 · 0浏览

[论文] Masked Visual Actions for Unified World Modeling

论文概要

研究领域: CV 作者: Hadi Alzayer, Wenlong Huang, Haonan Chen 发布时间: 2026-07-22 arXiv: 2507.17088

中文摘要

视频模型吸收了关于视觉世界如何运动、交互和对接触作出响应的丰富先验,使其成为机器人世界建模的有前景的基础。核心挑战在于如何以与模型学习这些交互先验的视觉空间相一致的形式,同时仍基于物理操作,向此类模型传达动作。我们引入\textbf{掩码视觉动作}(Masked Visual Actions),这是一种像素空间控制接口,将动作表示为视频中任意实体的部分揭示轨迹。揭示机器人运动使模型充当前向动力学模型,预测场景对低级机器人动作的响应;而揭示期望的物体运动则使同一模型恢复与该结果一致的机器人行为。仅使用15小时来自真实视频和仿真的掩码示例进行微调,单个检查点即可在多样化场景和多种具身形式中实现强大的视觉保真度和可控性。在下游操作任务中,该模型生成的想象推演结果与真实世界执行相关联,用于策略评估;通过基于模型的规划对候选未来进行排序来改进决策;并支持通过从期望物体运动合成机器人运动来进行逆向建模。

原文摘要

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of m...

--- *自动采集于 2026-07-23*

#论文 #arXiv #CV #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens