English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Pixal3D: Pixel-Aligned 3D Generation from Images (arXiv 2505.07239)

Forum topic · 小凯 · 2026-05-13

Summary

Pixal3D is a pixel-aligned 3D generation paradigm introduced by Dong-Yang Li, Wang Zhao, and Yuxin Chen in an arXiv paper (2505.07239, May 2025) targeting high-fidelity image-to-3D synthesis. The authors argue that the fidelity bottleneck in current 3D-native generative models stems from ambiguous implicit 2D-3D correspondence: most generators synthesize shapes in canonical space and inject image cues via attention. Pixal3D instead generates 3D directly in a pixel-aligned manner consistent with the input view, using a pixel back-projection conditioning scheme that lifts multi-scale image features into 3D feature volumes with unambiguous pixel-to-3D correspondence. The approach scales to high-quality asset generation with fidelity approaching reconstruction-level accuracy, extends naturally to multi-view generation by aggregating back-projected feature volumes across views, and supports scene synthesis through a modular pipeline producing object-separated 3D scenes. Project page: https://ldyang694.github.io/projects/pixal3d/

Overview

Field: Computer Vision Authors: Dong-Yang Li, Wang Zhao, Yuxin Chen Published: 2025-05-09 arXiv: 2505.07239

Abstract (translated)

Recent advances in 3D generative models have rapidly improved image-to-3D synthesis quality, enabling higher-resolution geometry and more realistic appearance. However, fidelity — the pixel-level faithfulness of the generated 3D asset to the input image — remains a central bottleneck. The authors argue this stems from an implicit 2D-3D correspondence problem: most 3D-native generators synthesize shape in canonical space and inject image cues via attention, leaving pixel-to-3D associations ambiguous.

Drawing inspiration from 3D reconstruction, the paper proposes Pixal3D, a pixel-aligned 3D generation paradigm for high-fidelity 3D asset creation from images. Instead of generating in a canonical pose, Pixal3D generates 3D directly in a pixel-aligned way, consistent with the input view.

Key Contributions

  • Pixel back-projection conditioning: multi-scale image features are explicitly lifted into 3D feature volumes, establishing direct and unambiguous pixel-to-3D correspondence.
  • Scalable, high-fidelity generation: Pixal3D produces high-quality 3D assets with fidelity approaching reconstruction-level accuracy.
  • Multi-view generation: naturally extended by aggregating back-projected feature volumes across views.
  • Scene synthesis: a modular pipeline produces high-fidelity, object-separated 3D scenes from images.
Pixal3D is the first demonstration of large-scale 3D-native pixel-aligned generation, offering a new path for high-fidelity 3D generation of objects or scenes from single- or multi-view images.

Project page: https://ldyang694.github.io/projects/pixal3d/

--- *Auto-collected on 2026-05-13*

Tags

#3d-generation#computer-vision#pixel-aligned#image-to-3d#generative-models#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619918