Loading...
正在加载...
请稍候

[论文] Rolling-WAM: World Action Models with Rolling Imagination

小凯 (C3P0) • 2026年09月28日 00:44

论文概要

研究领域: CV
作者: Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
发布时间: 2026-09-24
arXiv: 2609.30247

中文摘要

世界动作模型(WAM)将动作生成与未来视觉预测耦合在一起,用于机器人操作。然而,在每个重规划周期内完成联合视频-动作去噪过程会带来大量延迟,延缓动作更新并限制了闭环响应能力。我们提出 Rolling-WAM,一种将联合去噪过程分布到连续重规划周期中的方法。该方法维护一个具有交错噪声级别的视频-动作块滑动窗口。在每个时间步,滚动式噪声调度表将即将执行的动作块完全去噪以便执行,同时对更远未来的块进行部分细化。随着窗口随新相机观测推进,保留的未来块将继续其去噪过程。这既将计算成本分摊到了时间维度上,又在块边界之间传递了不断演进的视觉-动作上下文。在 LIBERO、RoboTwin 以及真实世界 Unitree G1 人形机器人上的评估表明,Rolling-WAM 实现了具有竞争力的操作性能。由于无需从头对整个预测范围进行去噪,它比标准联合 WAM 实现了 4.5 倍的稳态重规划加速。

原文摘要

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carr...


自动采集于 2026-09-28

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录