English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CLAP: Cross-Embodiment Video World Models for Zero-Shot Physical Simulation

Forum topic · 小凯 · 2026-08-30

Summary

CLAP is a cross-embodiment action-conditioned video generation framework introduced by Kechen Liu and Ola Shorinwa (arXiv:2608.27406). State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from exploiting the vast corpus of heterogeneous video data containing rich signals for learning generalizable physics. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor, enabling training on diverse, internet-scale videos spanning both human and robotic agents. To handle sharply differing action representations across robot platforms and their absence in human videos, CLAP unifies action spaces using end-effector poses, language instructions, and latent actions. It further introduces a curriculum-based cross-embodiment learning recipe: it first learns foundational physics priors from unlabeled videos via latent actions, then anchors them in the end-effector action space for zero-shot deployment on real-world tasks. CLAP approaches or exceeds the performance of state-of-the-art single-embodiment video models in challenging environments such as DROID.

Paper Overview

Field: Computer Vision Authors: Kechen Liu, Ola Shorinwa Published: 2026-08-27 arXiv: 2608.27406

Abstract

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, the authors introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents.

CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos.

Key Contributions

1. Unified action space: CLAP harmonizes different action representations using end-effector poses, language instructions, and latent actions. 2. Curriculum-based cross-embodiment learning: To address the limitations of each representation, CLAP first learns foundational physics priors from unlabeled video data via latent actions, then anchors these priors in the end-effector action space for zero-shot deployment to real-world tasks.

Results

CLAP approaches or exceeds the performance of state-of-the-art single-embodiment video models in challenging environments such as DROID.

---

*Auto-collected on 2026-08-30.*

Tags

#paper#arxiv#computer-vision#video-generation#world-models#robotics#cross-embodiment

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634243