Summary
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement in robotics, but their effectiveness rests on an unverified assumption: that generated futures faithfully reflect arbitrary valid actions. This paper introduces WorldEcho, a diagnostic framework that probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. The diagnosis reveals that current world models execute expert actions reasonably well but struggle with diverse off-expert trajectories, either ignoring commanded actions or producing visually invalid rollouts. To address these failures, the authors propose WorldSync, which strengthens action following along three complementary dimensions: distribution coverage, representation grounding, and intervention-effect alignment. WorldSync expands the training distribution of action consequences, uses action forcing to ground intermediate video representations in action-induced robot dynamics, and aligns predicted changes under action interventions with corresponding changes in real futures. Experiments on the RoboTwin benchmark and real-world robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to reach higher success rates. arXiv: 2608.24885.
Paper Overview
Field: Robotics
Authors: Sixiang Chen, Jiaming Liu, Jixian Wu, Yichen Guo, Tinghao Wang, et al.
arXiv: 2608.24885
Abstract
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated.
To address this gap, the authors introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. The diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories — either ignoring the commanded actions or producing visually invalid rollouts.
The authors further propose WorldSync, which strengthens action following along three complementary dimensions:
- Distribution coverage: expands the training distribution of action consequences.
- Representation grounding: uses action forcing to have experts ground intermediate video representations in action-induced robot dynamics.
- Intervention-effect alignment: aligns predicted changes under action interventions with the corresponding changes in real futures.
Experiments on the
RoboTwin benchmark and real-world robot tasks demonstrate that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to reach higher success rates.
*Auto-collected on 2026-08-27.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178634083