English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Fields

Forum topic · 小凯 · 2026-05-16

Summary

VGGT-Edit is a feed-forward framework for text-conditioned native 3D scene editing, proposed by Kaixin Zhu, Yiwen Tang, and Yifan Yang (arXiv:2505.08632). Existing 3D editing methods rely on 2D-lifting pipelines that edit each view independently and re-project into 3D, often causing blurry textures and inconsistent geometry. VGGT-Edit instead introduces depth-synchronized text injection that aligns semantic guidance with the backbone's spatial pose alignment for stable instruction grounding. A residual transformation head then directly predicts 3D geometric displacements to deform the scene while keeping the background stable. Training is supervised with a multi-project objective enforcing geometric accuracy and cross-view consistency. The authors also release DeltaScene, a large-scale dataset generated via an automated pipeline with 3D-consistency filtering. Experiments show VGGT-Edit outperforms 2D-lifting baselines with sharper object details, stronger multi-view consistency, and near-instant inference.

Overview

  • Field: Computer Vision
  • Authors: Kaixin Zhu, Yiwen Tang, Yifan Yang
  • Published: 2026-05-16
  • arXiv: 2505.08632
  • Abstract (translation)

    High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environments in a single forward pass. However, despite their strong performance in static scene perception, these models remain limited in responding to dynamic human instructions, which restricts their use in interactive applications.

    Existing editing methods typically rely on a 2D-lifting strategy, where individual views are edited independently and then lifted back into 3D space. This indirect pipeline often leads to blurry textures and inconsistent geometry, as 2D editors lack the spatial awareness required to preserve structure across viewpoints.

    Contributions

    To address these limitations, the authors propose VGGT-Edit, a feed-forward framework for text-conditioned native 3D scene editing:

  • Depth-synchronized text injection: aligns semantic guidance with the backbone's spatial pose alignment, ensuring stable instruction grounding.
  • Residual transformation head: directly predicts 3D geometric displacements to deform the scene while keeping the background stable.
  • Multi-project objective: supervises the framework to enforce geometric precision and cross-view consistency.
  • DeltaScene dataset: a large-scale dataset generated via an automated pipeline with 3D-consistency filtering to ensure ground-truth quality.

Results

Experiments show that VGGT-Edit significantly outperforms 2D-lifting baselines, producing sharper object details, stronger multi-view consistency, and near-instant inference speed.

*Auto-collected on 2026-05-16.*

Tags

#3d-scene-editing#feed-forward#text-conditioned#vggt#computer-vision#deltascene#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620089