English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

Forum topic · 小凯 · 2026-07-08

Summary

PixWorld is a single diffusion-based model that unifies 3D scene generation and reconstruction in pixel space. Traditionally, 3D reconstruction relies on pixel-based regression while generation uses latent diffusion; prior attempts to merge them in latent space suffer from information loss from latent encoding, diffusion objectives detached from the underlying 3D representation, and dependence on pre-trained VAEs or RAEs. PixWorld instead supervises diffusion directly on rendered images, aligning optimization with 3D scene fidelity and eliminating these limitations. Beyond photometric and perceptual losses that lack 3D geometric awareness, it introduces a geometry-aware loss that aligns rendered views with ground truth in the geometry-aware feature space of a pre-trained 3D foundation model, providing explicit 3D structural supervision. According to the authors (Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian), PixWorld consistently outperforms prior latent-space generative methods and matches state-of-the-art reconstruction approaches. Paper: arXiv 2607.05373.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian
  • Published: 2026-07-06
  • arXiv: 2607.05373

Abstract (translated)

3D reconstruction and generation are typically handled by different paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works have attempted to unify them in latent space, but this approach has notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, both branches suffer information loss introduced by latent encoding, and a pre-trained VAE or RAE is required.

This work re-unifies the two tasks under a pixel-space diffusion paradigm, proposing PixWorld — a single model that jointly handles 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld eliminates the above limitations and aligns optimization with 3D scene fidelity.

In addition to photometric and perceptual supervision, which operate only at the 2D image level and lack 3D geometric awareness, the authors introduce a geometry-aware loss that aligns rendered views with ground truth in the geometry-aware feature space of a pre-trained 3D foundation model, providing 3D structural supervision.

Results

PixWorld consistently outperforms previous latent-space generative methods and matches state-of-the-art reconstruction methods.

---

*Source: zhichai.net forum post, auto-collected 2026-07-06.*

Tags

#3d-reconstruction#3d-generation#diffusion-models#computer-vision#pixel-space#geometry-aware-loss#arxiv#pixworld

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346225