English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Vista4D: Video Reshooting with 4D Point Clouds

Forum topic · 小凯 · 2026-04-25

Summary

Vista4D is a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud, enabling re-synthesis of a dynamic scene from new camera trajectories and viewpoints. Existing video reshooting methods often struggle with depth estimation artifacts in real-world dynamic videos, fail to preserve content appearance, and lack precise camera control for challenging trajectories. Vista4D addresses this with a 4D-grounded point cloud representation built via static pixel segmentation and 4D reconstruction, which explicitly preserves seen content and provides rich camera signals. The model is trained on reconstructed multiview dynamic data to improve robustness against point cloud artifacts during real-world inference. Results show improved 4D consistency, camera control, and visual quality compared to state-of-the-art baselines across various videos and camera paths. The method also generalizes to real-world applications such as dynamic scene extension and 4D scene re-composition. Paper: arXiv 2604.21938.

Paper Overview

Field: Computer Vision (CV) Authors: Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca Published: 2026-04-23 arXiv: 2604.21938

What Vista4D Does

Vista4D is a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Given an input video, the method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint.

Motivation

Existing video reshooting methods typically suffer from:

  • Depth estimation artifacts in real-world dynamic videos
  • Failure to preserve content appearance
  • Imprecise camera control for challenging new trajectories
  • Method

  • Build a 4D-grounded point cloud representation using static pixel segmentation and 4D reconstruction
  • Explicitly preserve seen content while providing rich camera signals
  • Train with reconstructed multiview dynamic data for robustness against point cloud artifacts during real-world inference
  • Results and Applications

    Compared with state-of-the-art baselines across various videos and camera paths, Vista4D shows improvements in:

  • 4D consistency
  • Camera control
  • Visual quality
The method also generalizes to real-world applications such as dynamic scene extension and 4D scene re-composition.

Abstract (Original)

> We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing video reshooting methods often struggle with depth estimation artifacts of real-world dynamic videos, while also failing to preserve content appearance and failing to maintain precise camera control for challenging new trajectories. We build a 4D-grounded point cloud representation with static pixel segmentation and 4D reconstruction to explicitly preserve seen content and provide rich camera signals, and we train with reconstructed multiview dynamic data for robustness against point cloud artifacts dur...

---

*Auto-collected on 2026-04-25*

Tags

#vista4d#computer-vision#video-reshooting#4d-point-clouds#camera-control#4d-reconstruction#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618736