Summary
iMaC (Image as Action Control) is a new unified control paradigm for embodied world models that uses raw visual images as native action representations instead of low-dimensional structured action vectors such as joint angles or end-effector poses. Traditional embodied frameworks suffer from limited expressiveness, poor generalization across different robot embodiments, and unnatural dynamic modeling of complex physical interactions. iMaC formulates continuous visual manipulation as image-based action tokens that inherently capture spatial motion intent, interaction geometry constraints, and subtle physical dynamics. The system uses a dual-branch embodied architecture: an image action encoder that compresses goal-driven visual images into compact action embeddings, and a dynamics world predictor that learns environment transition rules conditioned on image actions for high-fidelity future state prediction and closed-loop embodied control. Experiments on public embodied manipulation benchmarks and real-world robot scenarios show that iMaC outperforms vector-based action control baselines in prediction accuracy, task success rate, and cross-scene generalization. Its image-based action design also removes the dependency on manually defined action spaces, enabling flexible universal control of heterogeneous embodied agents. Paper: arXiv:2506.04834.
Overview
Research area: Computer Vision (CV)
Authors: Zhenyu Wu, Xiuwei Xu, Yukun Zhou
Published: 2025-06-06
arXiv: 2506.04834
Summary
Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely on low-dimensional structured action vectors (e.g., joint angles and end-effector poses), which suffer from limited expressive capacity, poor generalization across diverse embodiments, and unnatural dynamic modeling for complex physical interactions.
To address these limitations, this paper proposes iMaC (Image as Action Control), a novel unified control paradigm that treats raw visual images as native action representations for embodied world models. Departing from traditional explicit kinematic action encoding, iMaC formulates continuous visual manipulation as image-based action tokens, which inherently encapsulate spatial motion intent, interaction geometry constraints, and subtle physical dynamics.
Architecture
The authors build a dual-branch embodied architecture consisting of:
- Image action encoder: compresses goal-driven visual images into compact action embeddings.
- Dynamics world predictor: learns environment transition rules conditioned on image actions, enabling high-fidelity future state prediction and closed-loop embodied control.
Results
Extensive experiments on public embodied manipulation benchmarks and real-world robotic scenarios demonstrate that iMaC outperforms vector-based action control baselines in:
- Prediction accuracy
- Task success rate
- Cross-scene generalization
Additionally, the image-based action design eliminates reliance on manually defined action spaces, enabling flexible and universal control across heterogeneous embodied agents.
---
Source: arXiv:2506.04834, auto-collected forum post.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177981043