Summary
This paper investigates the reliability of multiview 3D consistency metrics used to evaluate novel view synthesis (NVS) and sparse-view reconstruction. Standard evaluation assumes images depict a single static scene, but real outputs often contain artifacts, outlier frames, repeated views, or noise that still receive high scores. The authors systematically compare neural reconstruction priors with classical geometric verification, introducing a controlled robustness benchmark for multiview consistency and a parametric framework that decomposes neural metrics into backbone, residual, and aggregation components. This framework reproduces MEt3R and yields variants with up to 3x improved robustness. Analysis reveals that VGGT, MASt3R, DUSt3R, and Fast3R hallucinate dense geometry and cross-view support for unrelated scenes, duplicate images, and random noise. The authors further propose COLMAP-based failure-aware metrics using matching, registration, dense support, and reconstruction failure as consistency signals, achieving up to 4x higher correlation with human judgment than MEt3R on real NVS outputs.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620479