English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SCULPT: Subtractive Composition for 3D Part Generation

Forum topic · 小凯 · 2026-08-15

Summary

SCULPT is a part-aware 3D generation framework that creates complete digital assets while exposing structural parts for editing, material assignment, animation, and reuse. Unlike segmentation-based methods that partition an already generated shape, or additive methods that assemble parts from predefined layouts, SCULPT uses subtractive composition: starting from a complete object in a structured 3D latent space, it iteratively applies a joint split predictor to generate one extracted part together with the remaining object. The predictor performs a coupled denoising process conditioned on both the input image and the current 3D state, so parts and the remainder are generated together rather than reconciled afterward. Operating on the union of native sparse 3D supports, neighboring parts may overlap instead of being forced into a disjoint voxel partition. The rollout stops when the remainder support becomes empty or a fixed safety cap is reached, letting part count adapt per object. Experiments show state-of-the-art geometry on PartObjaverse, strong complete-object reconstruction after assembly, and fine-grained textured part decomposition on dataset images, a text-to-image input, and a real photograph. Paper: arXiv 2608.13541.

Paper Overview

Field: Computer Vision Authors: Sikuang Li, Chen Yang, Jiemin Fang, Jiazhong Cen, Yuhe Wei, Jichen Pang, Wei Shen, Qi Tian Published: 2026-08-13 arXiv: 2608.13541

Abstract

Part-aware 3D generation aims to create digital assets that are coherent as complete objects while exposing structural parts for editing, material assignment, animation, and reuse. Existing methods impose this structure outside the native generation loop: segmentation-based methods partition an already generated shape, while additive methods synthesize parts from predefined layouts, boxes, or tokens and then reconcile them into a whole. The former preserves the generated geometry but fixes the object before part boundaries are determined; the latter exposes part cardinality but often leaves shared boundaries vulnerable to gaps, interpenetrations, and material discontinuities.

In this paper, the authors propose SCULPT, a framework that addresses these challenges through subtractive composition. Given a complete object represented in a structured 3D latent space, SCULPT iteratively applies a joint split predictor to generate one extracted part together with the remaining object. The predictor performs a coupled denoising process conditioned on both the image and the current 3D state, so the extracted part and updated remainder are generated together rather than reconciled after generation. The joint split predictor processes both outputs on the union of their native sparse 3D supports, allowing neighboring supports to overlap rather than imposing a disjoint voxel partition. The rollout ends when the remainder support becomes empty or reaches a fixed safety cap, allowing the number of generated parts to adapt to each object within that bound.

Key Results

  • State-of-the-art geometry performance on PartObjaverse.
  • Strong complete-object reconstruction preserved after part assembly.
  • Fine-grained textured part decomposition demonstrated on four dataset images, one text-to-image-generated input, and one real-world photograph, going beyond standard benchmarks.
---

*Auto-collected on 2026-08-15.*

Tags

#3d-generation#part-aware-generation#computer-vision#generative-models#diffusion#arxiv#3d-assets

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633503