Paper Overview
- Research field: Computer Vision (CV)
- Authors: Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schönberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, Christian Rupprecht
- arXiv: 2605.15195
- Demonstrates predictable scaling of reconstruction quality with model and data size
- Improves accuracy and efficiency on both static and dynamic scenes
- Introduces training-efficiency-oriented architectural changes and a single dense prediction head with multi-task supervision
- Provides a high-quality annotation pipeline for dynamic scenes and a self-supervised learning protocol
> Note: The Chinese summary included in the original post describes a different paper (an agent-memory framework called "Preping"), which appears to be a mismatch in the source. The original abstract below, which corresponds to the title and arXiv link, covers VGGT-Ω.
Original Abstract
Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these models scales predictably with model and data size. We do so by introducing VGGT-Ω, which substantially improves reconstruction accuracy, efficiency, and capabilities for both static and dynamic scenes. To enable training this model at an unprecedented scale, we introduce architectural changes that improve training efficiency, a high-quality data annotation pipeline that supports dynamic scenes, and a self-supervised learning protocol. We simplify VGGT's architecture by using a single dense prediction head with multi-task supervision and removing the ...
*(Abstract truncated in source; see the arXiv page for the full text.)*