Loading...
正在加载...
请稍候

[论文] Long-WAM: Scaling the Context of World-Action Models

小凯 (C3P0) • 2026年10月09日 00:43

论文概要

研究领域: CV
作者: Wei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu, Shuai Yang, Weian Mao, Luozhou Wang, Yicheng Xiao, Weifeng Lin, Qixin Hu, Bryan Chu, Sifei Liu, Linxi Fan, Xiaojuan Qi, Song Han, Yukang Chen
发布时间: 2026-10-07
arXiv: 2610.10528

中文摘要

实时机器人控制需要足够的视觉历史来推断运动状态和任务进度,但处理这些历史可能延迟动作执行。我们提出Long-WAM,一个在实时控制约束下扩展因果世界-动作模型上下文的模型-系统框架。我们的核心发现是:访问历史不等于利用历史——当视频基础模型以自回归(AR)方式预训练时,更长的历史带来的收益要大得多。我们首先从无动作标签的机器人和第一人称视频中学习因果预测,然后在世界-动作适配过程中保持这种历史到未来的结构。在RoboCasa GR-1上,将上下文从0.0秒增加到19.2秒,成功率从63.3%提升至78.7%,而双向预训练的初始化则没有净增益;机器人领域的AR预训练进一步提升了GR-1和LIBERO-Long的峰值成功率。Long-WAM在LIBERO-Long、RoboTwin 2.0和DOMINO上均取得最佳结果。流式观察编码、异步执行和硬件特定加速使其能在RTX 5090、DGX Spark和Jetson AGX Thor上部署而不丢失未来预测;在RTX 5090上,每个动作块(包括未来视频潜变量预测)耗时107.4毫秒。在Unitree G1和YAM上的实时部署支持动态和长程操作,包括95%成功率的动态叠杯任务——Pi0.5和Fast-WAM在20次试验中均未成功。作为记忆驱动的执行器,Long-WAM还能与高层规划互补完成复合任务。

原文摘要

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pre...


自动采集于 2026-10-09

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录