[论文] LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC

研究领域: CV 作者: Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy 发布时间: 2026-10-08 arXiv: 2610.12407

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy 发布时间: 2026-10-08 arXiv: 2610.12407

中文摘要

世界动作模型(WAM)预测动作和未来观测,通常基于重建式表征,这种表征携带噪声和冗余信息,可能使下游预测复杂化。我们引入 LeWAM——一个用于前向、后向、逆动力学和策略预测的双向 Transformer,基于通过全部四种模式端到端训练的无解码器 JEPA 隐空间。我们看到以下收益:1)对齐:线性探针从 LeWAM 的隐空间中读取机器人和物体状态的效果优于从常规 Le World Model(仅前向的 JEPA 世界模型)中读取,同时该隐空间对视觉干扰物的忽略能力与 LeWM 相当,远优于基于重建的 WAM。2)行动:LeWAM 的闭环评估与同编码器上训练的常规流匹配策略在同等规模下表现相当,同时额外提供了一个世界模型。3)规划:用 WAM 规划时采样原始动作会让 MPC 利用动力学模型的不准确之处;改为在策略头的噪声空间中规划可以改善这些 WAM 的闭环性能。

原文摘要

World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at...


*自动采集于 2026-10-11*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens