English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

Forum topic · 小凯 · 2026-06-03

Summary

RoboDream (arXiv:2506.00003) is a research paper introducing a generalizable, embodiment-centric world model for scalable robot data synthesis. Real-world robot learning requires large-scale, diverse demonstrations, but teleoperated data collection is expensive and slow. Existing video diffusion approaches often produce only superficial visual augmentation or suffer from embodiment hallucinations with physically infeasible motions. RoboDream anchors generation to rendered robot motion while conditioning on explicit scene and object priors, decoupling trajectory execution from environment synthesis. This enables photorealistic demonstrations with novel objects, scenes, and viewpoints. The framework supports two data-scaling capabilities: retrieval and rebirth (reusing existing trajectories in entirely new contexts without new motion data) and prop-free teleoperation (operators gesture in the air while the model hallucinates the target objects and scene, eliminating reset time). Real-world experiments show generated data consistently improves downstream policy performance and substantially reduces real-world data requirements across manipulation tasks.

Paper Overview

  • Field: Computer Vision / Robotics
  • Authors: Junjie Ye, Rong Xue, Basile Van Hoorick
  • Published: 2026-06-03
  • arXiv: 2506.00003

Abstract

Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.

RoboDream is a generalizable, embodiment-centric world model that achieves scalable data generation by synthesizing photorealistic demonstrations with novel objects, in novel scenes, and from novel viewpoints. The approach anchors generation to rendered robot motion while conditioning on explicit scene and object priors, effectively decoupling trajectory execution from environment synthesis.

Key Capabilities

This formulation unlocks two powerful data-scaling capabilities:

1. Retrieval and rebirth — repurposing existing trajectories for entirely new contexts without any new motion data. 2. Prop-free teleoperation — the operator manipulates in the air, and the model then hallucinates the target objects and scene, eliminating reset time.

Results

Through real-world experiments, the authors demonstrate that the generated data consistently improves downstream policy performance and significantly reduces real-world data requirements across a variety of manipulation tasks.

---

*Auto-collected on 2026-06-03*

Tags

#robotics#world-models#video-diffusion#data-synthesis#manipulation#paper#arxiv#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980768