论文概要
研究领域: ML
作者: Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan, Sarath Chandar, Tushar Nagarajan, Daniel Severo, Koustuv Sinha, Michal Drozdzal, Adriana Romero Soriano, Jeannette Bohg, Nicolas Ballas, Mahmoud Assran
发布时间: 2026-10-07
arXiv: 2610.10515
中文摘要
潜在世界模型在预测未来状态和真实世界规划方面展现出卓越能力。然而实践中,我们缺乏一种原则性的方法来估计其能力如何随模型规模、数据和计算量扩展——这个开放问题阻碍了领域进展。本工作提出RoboJEPA——基于联合嵌入预测架构(JEPA)的世界模型,在覆盖12种机器人形态的大规模数据集上训练。我们证明RoboJEPA的想象误差(其潜在rollout的误差)遵循关于计算量的二阶幂律,使我们能够预测远超拟合范围的模型质量。我们进一步表明,下游机器人规划性能随计算量可预测地改善,且想象误差与其强相关,使其成为真实机器人评估的可靠代理。最后,我们证明潜在世界模型可以零样本部署为机器人智能体,通过规划朝向单一目标图像来完成需要长程规划的真实硬件任务。我们发布了所有模型检查点以及训练和机器人部署代码。据我们所知,这是首个为多种形态机器人世界模型在真实机器人数据上建立扩展定律的工作,RoboJEPA以8B参数成为迄今最大的JEPA预测模型。
原文摘要
Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit. We further show that downstream robotic planning performance improves predictably with compute, and that imagination error is strongly...
自动采集于 2026-10-09
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。