[论文] WorldSculpt: Generating Compositional Worlds from Grounded Videos

研究领域: CV 作者: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang…

论文概要

研究领域: CV 作者: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang 发布时间: 2026-09-04 arXiv: 2609.05416

中文摘要

翻译缺失

原文摘要

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated com...


*自动采集于 2026-09-08*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(1)

Q

难得,这帖的 arXiv 号是真的(2609.05416),12 位作者一字不差:Alaya Lab(盛大系)加东京大学,代码 Apache-2.0 已开源。基准也有意思,UE-MeshyScene 六个场景,93 到 701 个物体不等,"数百物体"没有吹牛。听着很美,下面泼两瓢冷水。

头一瓢:标题里最重的词是 Grounded,而"接地"其实是买来的。相机位姿、实例 mask(SAM3 跟踪)、3D 定位框,全是假定已经拿到的输入,论文明说这一步超出本文范围。名字读起来像"从生肉视频里自己摸清世界",工序上更像"别人切好菜,我来炒"。

第二瓢:SOTA 先按住,别急着鼓掌。旗舰基准上只比了 Meta 的 ShapeR 一个基线,中位数误差好出一个量级是真的,"比无可比"也是真的。另外只出几何不出纹理,只认静态场景,离机器人能跑的仿真还差好几道坎。

它真正讨喜的位置在别处。如今的世界模型(Marble 这类)端出来的是一团没分化的高斯汤:好看,没法编辑,你没法指着说"这朵是花瓶、那只是猫"。把汤倒回一副副独立 mesh、各就各位送进引擎,补的是世界模型通向游戏和仿真的最后一截管道。至于遮挡区的洞是"编"出来的——编得能拆能搬,就有用。

暂无表态

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens