Summary
VGGT-Edit is a feed-forward framework for text-conditioned native 3D scene editing, introduced by Kaixin Zhu and colleagues in an arXiv paper (2605.15186). While generalizable feed-forward architectures can now reconstruct complex 3D environments in a single forward pass, they respond poorly to dynamic human instructions. Existing editing pipelines lift independently edited 2D views back into 3D, often causing blurry textures and inconsistent geometry because 2D editors lack spatial awareness across viewpoints. VGGT-Edit addresses this with depth-synchronized text injection, aligning semantic guidance with the backbone's spatial pose alignment for stable instruction grounding, and a residual transformation head that directly predicts 3D geometric displacements to deform scenes while keeping the background stable. Training uses a multi-term objective enforcing geometric accuracy and cross-view consistency. The authors also introduce DeltaScene, a large-scale dataset generated via an automatic pipeline with 3D-consistency filtering. Experiments show VGGT-Edit substantially outperforms 2D-lifting baselines, delivering sharper object details, stronger multi-view consistency, and near real-time inference.
Overview
Research area: Computer Vision
Authors: Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang, Bohan Zeng, Ziyu Guo, Ruichuan An, Zhou Liu, Qizhi Chen, Delin Qu, Jaehong Yoon, Wentao Zhang
Published: 2026-05-14
arXiv: 2605.15186
Abstract
High-quality 3D scene reconstruction has recently advanced to generalizable feed-forward architectures capable of generating complex environments in a single forward pass. However, despite their strong static-scene perception, these models remain limited in responding to dynamic human instructions, restricting their use in interactive applications.
Existing editing methods typically rely on 2D-lifting strategies, where individual views are edited independently and then lifted back into 3D space. This indirect pipeline often results in blurry textures and inconsistent geometry, since 2D editors lack the spatial awareness required to preserve structure across viewpoints.
To address these limitations, the authors propose VGGT-Edit, a feed-forward framework for text-conditioned native 3D scene editing. Key contributions:
- Depth-synchronized text injection: aligns semantic guidance with the backbone's spatial pose alignment, ensuring stable instruction grounding.
- Residual transformation head: processes the semantic signal to directly predict 3D geometric displacements that deform the scene while keeping the background stable.
- Multi-term objective: supervises training to enforce geometric accuracy and cross-view consistency.
- DeltaScene dataset: a large-scale dataset generated via an automatic pipeline with 3D-consistency filtering to ensure realistic quality.
Experiments show that VGGT-Edit substantially outperforms 2D-lifting baselines, producing sharper object details, stronger multi-view consistency, and near-instant inference speed.
*Auto-collected on 2026-05-17.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620164