English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldSculpt: Generating Compositional 3D Worlds from Grounded Videos

Forum topic · 小凯 · 2026-09-08

Summary

WorldSculpt (arXiv:2609.05416) addresses the challenge of generating a compositional 3D representation of cluttered real-world scenes containing hundreds of objects. The method represents a scene as a collection of individual object meshes placed in a shared world frame, a format directly usable by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is difficult because densely packed objects heavily occlude one another, so each view reveals only a fraction of their geometry. Existing geometry-based approaches reconstruct the scene as a single representation, leaving incomplete geometry in occluded regions, while prior compositional methods with generative priors are limited to relatively simple scenes. WorldSculpt shows that complex, cluttered scenes with hundreds of objects can be generated compositionally by leveraging grounded videos. The paper is by Muyao Niu, Jixuan He, and colleagues including Zhixiang Wang, posted on arXiv on September 4, 2026, and categorized under computer vision.

Paper Overview

Research Area: Computer Vision (CV) Authors: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang Published: 2026-09-04 arXiv: 2609.05416

Abstract

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in obsented regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally...

*(Abstract truncated in the source post; see the arXiv page for the full text.)*

---

*Auto-collected on 2026-09-08.*

Tags

#worldsculpt#3d-generation#scene-reconstruction#compositional-3d#computer-vision#arxiv#object-meshes#generative-priors

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634616