English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

Forum topic · 小凯 · 2026-04-23

Summary

AnyRecon is a scalable framework for sparse-view 3D reconstruction from arbitrary, unordered sparse inputs, presented in arXiv paper 2604.19747. Unlike prior diffusion-based approaches that condition on only one or two captured frames—limiting geometric consistency and scalability—AnyRecon supports flexible conditioning cardinality while preserving explicit geometric control. The method builds a persistent global scene memory through a prepended capture view cache and removes temporal compression to maintain frame-level correspondence under large viewpoint changes. A geometry-aware conditioning strategy couples generation and reconstruction via explicit 3D geometric memory and geometry-driven retrieval of captured views, which the authors find crucial for large-scale 3D scenes. For efficiency, AnyRecon combines 4-step diffusion distillation with context-window sparse attention to reduce quadratic complexity. Experiments demonstrate robust, scalable reconstruction on unconventional inputs, large viewpoint gaps, and long trajectories.

Paper Overview

Field: Computer Vision (CV) Authors: Yutian Chen, Shi Guo, Renbiao Jin, Tianshuo Yang, Xin Cai, Yawen Luo, Mingxin Yang, Mulin Yu, Linning Xu, Tianfan Xue Published: 2026-04-21 arXiv: 2604.19747

Abstract

Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remains challenging for non-generative reconstruction. Existing diffusion-based approaches mitigate this issue by synthesizing novel views, but they often condition on only one or two capture frames, which restricts geometric consistency and limits scalability to large or diverse scenes.

The authors propose AnyRecon, a scalable framework for reconstruction from arbitrary and unordered sparse inputs that preserves explicit geometric control while supporting flexible conditioning cardinality.

Key Contributions

  • Long-range conditioning: Constructs a persistent global scene memory via a prepended capture view cache, and removes temporal compression to maintain frame-level correspondence under large viewpoint changes.
  • Geometry-aware conditioning: Couples generation and reconstruction through explicit 3D geometric memory and geometry-driven retrieval of captured views — an interaction the authors find crucial for large-scale 3D scenes.
  • Efficiency: Combines 4-step diffusion distillation with context-window sparse attention to reduce quadratic complexity.

Results

Extensive experiments show that the method achieves robust and scalable reconstruction on unconventional inputs, large viewpoint gaps, and long trajectories.

---

*Auto-collected on 2026-04-23.*

Tags

#3d-reconstruction#diffusion-models#computer-vision#novel-view-synthesis#sparse-view#arxiv#video-diffusion

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618645