论文概要
研究领域: CV
作者: Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun, Zequn Chen, Sen Wang, Xinjie Wang, Xia Chen, Jingfeng Yao, Weiheng Zhao, Shanglin Yuan, Zhizhong Su, Wei Sui, Wenyu Liu, Xinggang Wang
发布时间: 2026-09-29
arXiv: 2609.38163
中文摘要
世界-动作模型(World-action model)联合学习机器人策略和预测未来观测,使表示空间成为控制与预测之间的接口。我们通过受控对比研究该空间的设计,发现重建保真度和预训练感知特征都不能单独保证有效的策略学习。这些发现推动了ReWAM——一个以表示为中心的世界-动作模型,构建于预训练DINO特征之上。特征校准和时间表示瓶颈将这些特征组织为适合动力学建模的紧凑世界状态。动作锚定表示整形仅将动作损失梯度路由到瓶颈,从而让策略塑造表示编码的内容,同时世界模型学习其演化方式。无需生成视频预训练,ReWAM在RoboTwin 2.0上实现了93.6%的成功率。在RoboDojo上,使用约600小时的具身预训练数据,实现了平均12.29分和8.28%的成功率。
原文摘要
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. W...
自动采集于 2026-10-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。