English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Pixal3D: Pixel-Aligned 3D Generation from Images (arXiv 2505.07239)

Forum topic · 小凯 · 2026-05-13

Summary

Pixal3D is a pixel-aligned 3D generation paradigm proposed to address the fidelity bottleneck in image-to-3D synthesis. The authors argue that existing 3D-native generators suffer from an implicit 2D-3D correspondence problem: they synthesize shapes in canonical space and inject image cues via attention, leaving pixel-to-3D associations ambiguous. Inspired by 3D reconstruction, Pixal3D generates 3D directly in a pixel-aligned manner consistent with the input view. It introduces a pixel back-projection conditioning scheme that explicitly lifts multi-scale image features into 3D feature volumes, establishing unambiguous pixel-to-3D correspondences. The approach is scalable, produces high-quality 3D assets, and achieves fidelity approaching reconstruction-level accuracy. Pixal3D also naturally extends to multi-view generation by aggregating back-projected feature volumes across views, and supports scene synthesis through a modular pipeline producing high-fidelity, object-separated 3D scenes from images. According to the authors, this is the first demonstration of large-scale 3D-native pixel-aligned generation, offering a new path for high-fidelity object or scene 3D generation from single- or multi-view images. Paper: arXiv 2505.07239; project page: https://ldyang694.github.io/projects/pixal3d/

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Dong-Yang Li, Wang Zhao, Yuxin Chen
  • Published: 2025-05-09
  • arXiv: 2505.07239
  • Abstract

    Recent advances in 3D generative models have rapidly improved image-to-3D synthesis quality, enabling higher-resolution geometry and more realistic appearance. Yet fidelity — the pixel-level faithfulness of the generated 3D asset to the input image — remains a central bottleneck. The authors argue this stems from an implicit 2D-3D correspondence issue: most 3D-native generators synthesize shape in canonical space and inject image cues via attention, leaving pixel-to-3D associations ambiguous.

    Key Contributions

  • Pixel-aligned generation paradigm: Instead of generating in a canonical pose, Pixal3D generates 3D directly in a pixel-aligned way, consistent with the input view.
  • Pixel back-projection conditioning: A conditioning scheme explicitly lifts multi-scale image features into 3D feature volumes, establishing direct, unambiguous pixel-to-3D correspondences.
  • High fidelity at scale: Pixal3D is scalable, produces high-quality 3D assets, and achieves fidelity approaching reconstruction-level accuracy.
  • Multi-view extension: By aggregating back-projected feature volumes across views, Pixal3D naturally extends to multi-view generation.
  • Scene synthesis: The paper demonstrates benefits of pixel-aligned generation for scenes, proposing a modular pipeline that produces high-fidelity, object-separated 3D scenes from images.
  • Significance

    Pixal3D is presented as the first demonstration of large-scale 3D-native pixel-aligned generation, offering a new, instructive path for high-fidelity 3D generation of objects or scenes from single- or multi-view images.

    Links

  • Paper: https://arxiv.org/abs/2505.07239
  • Project page: https://ldyang694.github.io/projects/pixal3d/

Tags

#3d-generation#computer-vision#image-to-3d#pixel-aligned#generative-models#arxiv#reconstruction

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619918