论文概要
研究领域: CV
作者: Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
发布时间: 2026-09-24
arXiv: 2609.30247
中文摘要
世界动作模型(WAM)将动作生成与未来视觉预测耦合在一起,用于机器人操作。然而,在每个重规划周期内完成联合视频-动作去噪过程会带来大量延迟,延缓动作更新并限制了闭环响应能力。我们提出 Rolling-WAM,一种将联合去噪过程分布到连续重规划周期中的方法。该方法维护一个具有交错噪声级别的视频-动作块滑动窗口。在每个时间步,滚动式噪声调度表将即将执行的动作块完全去噪以便执行,同时对更远未来的块进行部分细化。随着窗口随新相机观测推进,保留的未来块将继续其去噪过程。这既将计算成本分摊到了时间维度上,又在块边界之间传递了不断演进的视觉-动作上下文。在 LIBERO、RoboTwin 以及真实世界 Unitree G1 人形机器人上的评估表明,Rolling-WAM 实现了具有竞争力的操作性能。由于无需从头对整个预测范围进行去噪,它比标准联合 WAM 实现了 4.5 倍的稳态重规划加速。
原文摘要
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carr...
自动采集于 2026-09-28
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。