English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

Forum topic · 小凯 · 2026-08-27

Summary

Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement in robotics, but their effectiveness rests on an unverified assumption: that generated futures faithfully reflect arbitrary valid actions. This paper introduces WorldEcho, a diagnostic framework that probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. The diagnosis reveals that current world models execute expert actions reasonably well but struggle with diverse off-expert trajectories, either ignoring commanded actions or producing visually invalid rollouts. To address these failures, the authors propose WorldSync, which strengthens action following along three complementary dimensions: distribution coverage, representation grounding, and intervention-effect alignment. WorldSync expands the training distribution of action consequences, uses action forcing to ground intermediate video representations in action-induced robot dynamics, and aligns predicted changes under action interventions with corresponding changes in real futures. Experiments on the RoboTwin benchmark and real-world robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to reach higher success rates. arXiv: 2608.24885.

Paper Overview

Field: Robotics Authors: Sixiang Chen, Jiaming Liu, Jixian Wu, Yichen Guo, Tinghao Wang, et al. arXiv: 2608.24885

Abstract

Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated.

To address this gap, the authors introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. The diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories — either ignoring the commanded actions or producing visually invalid rollouts.

The authors further propose WorldSync, which strengthens action following along three complementary dimensions:

  • Distribution coverage: expands the training distribution of action consequences.
  • Representation grounding: uses action forcing to have experts ground intermediate video representations in action-induced robot dynamics.
  • Intervention-effect alignment: aligns predicted changes under action interventions with the corresponding changes in real futures.
Experiments on the RoboTwin benchmark and real-world robot tasks demonstrate that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to reach higher success rates.

*Auto-collected on 2026-08-27.*

Tags

#robotics#world-models#policy-learning#action-conditioned-generation#simulation#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634083