Paper Overview
Field: Computer Vision Authors: Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant, Ryan Burgert, Yuancheng Xu, Koichi Namekata, Yiwei Zhao, Bolei Zhou, Micah Goldblum, Paul Debevec, Ning Yu Published: 2026-04-23 arXiv: 2604.21915
Abstract
We present Vista4D, a robust and flexible video reshooting framework that anchors the input video and target cameras in 4D point clouds. Specifically, given an input video, our method re-synthesizes the same dynamic scene from different camera trajectories and viewpoints. Existing video reshooting methods often struggle with depth estimation artifacts in real-world dynamic videos, fail to preserve content appearance, and cannot maintain precise camera control for challenging novel trajectories.
We construct an anchored 4D point cloud representation with static-pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals. We train with reconstructed multi-view dynamic data to be robust to point cloud artifacts during real-world inference. Our results show improvements in 4D consistency, camera control, and visual quality compared to state-of-the-art baselines across diverse videos and camera paths. Furthermore, our method generalizes to real-world applications such as dynamic scene extension and 4D scene re-composition.
--- *Automatically collected on 2026-04-27*