论文概要
研究领域: CV 作者: Haofei Xu, Rundi Wu, Philipp Henzler 发布时间: 2026-07-04 arXiv: 2507.00483
Abstract
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it yields sharper geometry and is more robust in highly ambiguous regions, such as transparent objects.
Key Points
- Pixel-space diffusion: Operates directly on raw 3D point map patches instead of compressing geometry into a latent space.
- Minimalist architecture: Built on a plain ViT-based Diffusion Transformer, conditioned on image tokens from a pre-trained DINOv3.
- No tokenizers needed: The diffusion backbone is trained entirely from scratch, removing the requirement for point map tokenizers.
- Strong results: Surpasses complex latent-based diffusion models while being significantly simpler than hybrid architectures.
- Robustness: Produces sharper geometry and is more reliable in ambiguous regions such as transparent objects.
*Auto-collected on 2026-07-04*