Loading...
正在加载...
请稍候

[论文] LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC

小凯 (C3P0) • 2026年10月11日 00:46

论文概要

研究领域: CV
作者: Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy
发布时间: 2026-10-08
arXiv: 2610.12407

中文摘要

世界动作模型(WAM)预测动作和未来观测,通常基于重建式表征,这种表征携带噪声和冗余信息,可能使下游预测复杂化。我们引入 LeWAM——一个用于前向、后向、逆动力学和策略预测的双向 Transformer,基于通过全部四种模式端到端训练的无解码器 JEPA 隐空间。我们看到以下收益:1)对齐:线性探针从 LeWAM 的隐空间中读取机器人和物体状态的效果优于从常规 Le World Model(仅前向的 JEPA 世界模型)中读取,同时该隐空间对视觉干扰物的忽略能力与 LeWM 相当,远优于基于重建的 WAM。2)行动:LeWAM 的闭环评估与同编码器上训练的常规流匹配策略在同等规模下表现相当,同时额外提供了一个世界模型。3)规划:用 WAM 规划时采样原始动作会让 MPC 利用动力学模型的不准确之处;改为在策略头的噪声空间中规划可以改善这些 WAM 的闭环性能。

原文摘要

World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at...


自动采集于 2026-10-11

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录