Paper Overview
Field: Computer Vision (CV) Authors: Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca Published: 2026-04-23 arXiv: 2604.21938
What Vista4D Does
Vista4D is a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Given an input video, the method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint.
Motivation
Existing video reshooting methods typically suffer from:
- Depth estimation artifacts in real-world dynamic videos
- Failure to preserve content appearance
- Imprecise camera control for challenging new trajectories
- Build a 4D-grounded point cloud representation using static pixel segmentation and 4D reconstruction
- Explicitly preserve seen content while providing rich camera signals
- Train with reconstructed multiview dynamic data for robustness against point cloud artifacts during real-world inference
- 4D consistency
- Camera control
- Visual quality
Method
Results and Applications
Compared with state-of-the-art baselines across various videos and camera paths, Vista4D shows improvements in:
Abstract (Original)
> We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing video reshooting methods often struggle with depth estimation artifacts of real-world dynamic videos, while also failing to preserve content appearance and failing to maintain precise camera control for challenging new trajectories. We build a 4D-grounded point cloud representation with static pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals, and we train with reconstructed multiview dynamic data for robustness against point cloud artifacts dur...
---
*Auto-collected on 2026-04-25*