English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible Surfaces

Forum topic · 小凯 · 2026-06-14

Summary

World Tracing is a generative pixel-aligned geometry representation for image-to-3D reconstruction that resolves the traditional trade-off between faithfulness and completeness. For each observed pixel, the model predicts an ordered sequence of 3D points: the first layer corresponds to the visible surface, while subsequent layers represent occluded surfaces behind it. This layered formulation allows the method to stay aligned with observed pixels while simultaneously completing geometry beyond what is visible in the input image. The work is by Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang, Yi Hua, Ben Mildenhall, Christoph Lassner, Narendra Ahuja, and Gengshan Yang, and was released on arXiv (2606.13652) in June 2026. The approach achieves strong performance across object-level, scene-level, and dynamic benchmarks, suggesting a unified framework for faithful and complete 3D prediction from images.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang, Yi Hua, Ben Mildenhall, Christoph Lassner, Narendra Ahuja, Gengshan Yang
  • Published: 2026-06-11
  • arXiv: 2606.13652
  • Abstract

    Image-to-3D methods trade off faithfulness and completeness. The authors introduce World Tracing, a generative pixel-aligned geometry representation that predicts 3D points aligned with observed pixels while completing geometry beyond visible surfaces.

    For each pixel, the model predicts ordered 3D points, where the first layer is the visible surface and subsequent layers represent occluded surfaces. The approach achieves strong performance on object, scene, and dynamic benchmarks.

    Key Idea

  • Pixel-aligned + generative: Predictions stay tied to observed pixels (faithful) while generating plausible geometry in occluded regions (complete).
  • Ordered layered points: Each pixel maps to multiple 3D points sorted by depth, with layer 1 = visible surface and deeper layers = occluded surfaces.
  • Broad evaluation: Validated on object-level, scene-level, and dynamic (video) benchmarks.
---

*Auto-collected on 2026-06-14.*

Tags

#computer-vision#3d-reconstruction#world-tracing#arxiv#generative-models#pixel-aligned-geometry#occlusion-completion

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981285