English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Action Images: End-to-End Policy Learning via Multiview Video Generation

Forum topic · 小凯 · 2026-04-09

Summary

This post introduces Action Images (arXiv:2504.06262), a unified world action model (WAM) that formulates robot policy learning as multiview video generation. Instead of encoding control as low-dimensional tokens or relying on separate action modules, the method converts 7-DoF robot actions into interpretable action images: 2D pixel-grounded multi-view action videos that explicitly track robot-arm motion. This pixel-grounded representation allows the video backbone itself to act as a zero-shot policy, eliminating the need for a separate policy head or action module and better exploiting pretrained video knowledge. Evaluated on RLBench and real-world tasks, the model achieves the strongest zero-shot success rates and outperforms prior video-space world models in joint video-action generation quality, suggesting interpretable action images are a promising direction for policy learning and cross-viewpoint, cross-environment transfer.

Paper Overview

Field: Computer Vision (CV) Authors: Haoyu Zhen, Zixian Gao, Qiao Sun Published: 2025-04-08 arXiv: 2504.06262

Introduction

World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model future states. However, existing approaches often rely on separate action modules, or use action representations that are not pixel-grounded, making it difficult to fully exploit the pretrained knowledge of video models and limiting transfer across viewpoints and environments.

Method

This work presents Action Images, a unified world action model that formulates policy learning as multiview video generation. Key ideas:

  • Instead of encoding control as low-dimensional tokens, 7-DoF robot actions are translated into interpretable action images: multi-view action videos grounded in 2D pixels that explicitly track robot-arm motion.
  • Because actions are represented directly in pixel space, the video backbone itself can serve as a zero-shot policy, with no separate policy head or action module required.
  • Results

  • On RLBench and real-world evaluations, the model achieves the strongest zero-shot success rates.
  • It outperforms prior video-space world models in joint video-action generation quality.

Conclusion

Interpretable action images offer a promising path for policy learning, enabling better use of pretrained video knowledge and improved transfer across viewpoints and environments.

--- *Auto-collected on 2026-04-09*

Tags

#robotics#computer-vision#world-models#policy-learning#video-generation#arxiv#embodied-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169676