English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Masked Visual Actions for Unified World Modeling

Forum topic · 小凯 · 2026-07-23

Summary

This paper introduces Masked Visual Actions (arXiv:2507.17088), a pixel-space control interface for robotic world modeling with video models. Action is expressed as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion lets the model act as a forward dynamics model predicting scene responses to low-level robot actions, while revealing desired object motion allows the same model to recover robot behavior consistent with that outcome. Finetuned with only about 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and embodiments. In downstream manipulation tasks, the model's imagined rollouts correlate with real-world execution for policy evaluation, improve decision-making via model-based planning that ranks candidate futures, and support inverse modeling by synthesizing robot motion from desired object motion.

Paper Overview

Field: Computer Vision / Robotics Authors: Hadi Alzayer, Wenlong Huang, Haonan Chen arXiv: 2507.17088

Abstract

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation.

The authors introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video:

  • Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions.
  • Revealing desired object motion makes the same model recover robot behavior consistent with that outcome (inverse modeling).
  • Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments.

    Key Capabilities

  • Policy evaluation: imagined rollouts generated by the model correlate with real-world execution.
  • Planning: model-based planning ranks candidate futures to improve decision-making.
  • Inverse modeling: synthesizes robot motion from desired object motion.
---

*Originally posted on zhichai.net, 2026-07-23.*

Tags

#robotics#world-modeling#computer-vision#video-models#manipulation#pixel-space-actions#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447023