English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MV-VDP: Multi-View Video Diffusion Policy for 3D Spatio-Temporal-Aware Robotic Manipulation

Forum topic · 小凯 · 2026-04-06

Summary

MV-VDP (Multi-View Video Diffusion Policy) is a new robot learning framework presented in arXiv paper 2604.03181 by Peiyan Li, Yixiang Chen, Yuan Xu, and colleagues. It addresses a key limitation of existing robotic manipulation policies: most rely on 2D visual observations and backbones pretrained on static image-text pairs, ignoring either the 3D spatial structure of the environment or its temporal evolution. This leads to high data requirements and poor understanding of environment dynamics. MV-VDP jointly models the 3D spatio-temporal state of the environment by simultaneously predicting multi-view heatmap videos and RGB videos. This design aligns the representation format of video pretraining with action finetuning, and specifies not only what actions the robot should take but also how the environment is expected to evolve in response to those actions. Experiments show MV-VDP can successfully perform complex real-world manipulation tasks using only ten demonstration trajectories.

Paper Overview

Field: Computer Vision (CV) / Robotics Authors: Peiyan Li, Yixiang Chen, Yuan Xu, et al. Published: 2026-04-03 arXiv: 2604.03181

Abstract

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained on static image-text pairs, resulting in high data requirements and limited understanding of environment dynamics.

To address this, the authors introduce MV-VDP, a multi-view video diffusion policy that jointly models the 3D spatio-temporal state of the environment. The core idea is to simultaneously predict multi-view heatmap videos and RGB videos, which:

1. Align the representation format of video pretraining with action finetuning, and 2. Specify not only what actions the robot should take, but also how the environment is expected to evolve in response to those actions.

Results

Extensive experiments demonstrate that MV-VDP can successfully execute complex real-world tasks using only ten demonstration trajectories, highlighting its strong sample efficiency.

--- *Auto-collected on 2026-04-06.*

Tags

#mv-vdp#video-diffusion-policy#robotic-manipulation#3d-spatio-temporal#multi-view#imitation-learning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169591