English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Vista4D: Video Reshooting with 4D Point Clouds

Forum topic · 小凯 · 2026-04-27

Summary

Vista4D is a robust and flexible video reshooting framework that anchors both the input video and target cameras in a 4D point cloud. Given an input video, the method re-synthesizes the same dynamic scene from different camera trajectories and viewpoints. Existing video reshooting approaches often struggle with depth estimation artifacts in real-world dynamic videos, fail to preserve content appearance, and cannot maintain precise camera control along challenging novel trajectories. Vista4D constructs an anchored 4D point cloud representation with static-pixel segmentation and 4D reconstruction, explicitly preserving seen content while providing rich camera signals. The model is trained on reconstructed multi-view dynamic data, making it robust to point cloud artifacts during real-world inference. Results show improved 4D consistency, camera control, and visual quality compared to state-of-the-art baselines across diverse videos and camera paths. The approach also generalizes to real-world applications such as dynamic scene extension and 4D scene re-composition.

Paper Overview

Field: Computer Vision Authors: Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant, Ryan Burgert, Yuancheng Xu, Koichi Namekata, Yiwei Zhao, Bolei Zhou, Micah Goldblum, Paul Debevec, Ning Yu Published: 2026-04-23 arXiv: 2604.21915

Abstract

We present Vista4D, a robust and flexible video reshooting framework that anchors the input video and target cameras in 4D point clouds. Specifically, given an input video, our method re-synthesizes the same dynamic scene from different camera trajectories and viewpoints. Existing video reshooting methods often struggle with depth estimation artifacts in real-world dynamic videos, fail to preserve content appearance, and cannot maintain precise camera control for challenging novel trajectories.

We construct an anchored 4D point cloud representation with static-pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals. We train with reconstructed multi-view dynamic data to be robust to point cloud artifacts during real-world inference. Our results show improvements in 4D consistency, camera control, and visual quality compared to state-of-the-art baselines across diverse videos and camera paths. Furthermore, our method generalizes to real-world applications such as dynamic scene extension and 4D scene re-composition.

--- *Automatically collected on 2026-04-27*

Tags

#video-reshooting#4d-point-clouds#computer-vision#novel-view-synthesis#camera-control#generative-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618799