[论文] CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simul...
研究领域: CV 作者: Kechen Liu, Ola Shorinwa 发布时间: 2026-08-27 arXiv: 2608.27406
论文概要
研究领域: CV 作者: Kechen Liu, Ola Shorinwa 发布时间: 2026-08-27 arXiv: 2608.27406
中文摘要
最先进的动作条件视频模型通常限于单一机器人形态,阻止它们利用包含丰富学习通用物理信号的异构视频数据语料库。为弥合这一差距,我们引入了CLAP,一个跨形态动作条件视频生成框架,能够训练于多样化、互联网规模的人类和机器人智能体视频。CLAP基于以下洞察:无论行为者如何,普遍物理定律支配时空动态。然而,跨形态学习并非易事,因为动作表示在不同机器人平台间差异很大,人类视频中通常缺失。CLAP通过以下核心贡献解决这一基本挑战。首先,CLAP使用末端执行器姿态、语言指令和潜在动作协调不同的动作空间。其次,为解决各自局限,CLAP引入了基于课程的跨形态学习配方,首先通过潜在动作在跨无标签视频数据中学习基础物理先验,然后将其锚定在末端执行器动作空间中以零样本部署到真实世界任务。CLAP接近或超越DROID等挑战性环境中的最先进单形态视频模型。
原文摘要
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contribution...
*自动采集于 2026-08-30*
#论文 #arXiv #CV #小凯