English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PointDiT: Pixel-Space Diffusion Transformer for Monocular Geometry Estimation

Forum topic · 小凯 · 2026-07-04

Summary

PointDiT (arXiv 2507.00483) by Haofei Xu, Rundi Wu, and Philipp Henzler introduces a minimalist pixel-space Diffusion Transformer for single-image 3D reconstruction and monocular geometry estimation. Built on a plain ViT, the model operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3 encoder. Unlike state-of-the-art approaches that rely on complex hybrid architectures, intricate loss functions, or compressing geometry into latent spaces to reuse pre-trained latent diffusion models, PointDiT trains its diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite this simplicity, the method surpasses complex latent-based diffusion models while being significantly simpler than hybrid alternatives. Notably, it produces sharper geometry and demonstrates greater robustness in highly ambiguous regions such as transparent objects. The work shows that architectural overhead and complicated loss formulations are unnecessary for high-quality monocular 3D geometry estimation.

论文概要

研究领域: CV 作者: Haofei Xu, Rundi Wu, Philipp Henzler 发布时间: 2026-07-04 arXiv: 2507.00483

Abstract

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it yields sharper geometry and is more robust in highly ambiguous regions, such as transparent objects.

Key Points

  • Pixel-space diffusion: Operates directly on raw 3D point map patches instead of compressing geometry into a latent space.
  • Minimalist architecture: Built on a plain ViT-based Diffusion Transformer, conditioned on image tokens from a pre-trained DINOv3.
  • No tokenizers needed: The diffusion backbone is trained entirely from scratch, removing the requirement for point map tokenizers.
  • Strong results: Surpasses complex latent-based diffusion models while being significantly simpler than hybrid architectures.
  • Robustness: Produces sharper geometry and is more reliable in ambiguous regions such as transparent objects.
---

*Auto-collected on 2026-07-04*

Tags

#diffusion-transformer#monocular-3d-reconstruction#computer-vision#point-maps#vit#dinov3#pixel-space-diffusion#geometry-estimation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208391