Paper Overview
Field: Computer Vision Authors: Kechen Liu, Ola Shorinwa Published: 2026-08-27 arXiv: 2608.27406
Abstract
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, the authors introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents.
CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos.
Key Contributions
1. Unified action space: CLAP harmonizes different action representations using end-effector poses, language instructions, and latent actions. 2. Curriculum-based cross-embodiment learning: To address the limitations of each representation, CLAP first learns foundational physics priors from unlabeled video data via latent actions, then anchors these priors in the end-effector action space for zero-shot deployment to real-world tasks.
Results
CLAP approaches or exceeds the performance of state-of-the-art single-embodiment video models in challenging environments such as DROID.
---
*Auto-collected on 2026-08-30.*