[论文] AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents
研究领域: CV 作者: Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, et al. 发布时间: 2026-09-30 arXiv: 2609.26761
论文概要
研究领域: CV 作者: Jiahao Zhang, Yeying Fan, Moitreya Chatterjee, et al. 发布时间: 2026-09-30 arXiv: 2609.26761
中文摘要
3D 组装任务需要将对零件及其关系的理解转化为精确的空间排列。预训练的通用智能体能否在无需额外组装特定微调的情况下,通过视觉交互完成物体组装?为研究这一问题,我们引入 AssemblyWorld——一个交互式 3D 环境,智能体在其中检查渲染视图并操作提供的刚性零件,在有条件时参考图像或组装手册。智能体通过 2D 视图感知零件几何形状,而非直接访问网格顶点或面,而其组装结果通过几何方式评估。在此环境基础上,我们构建了 AssemblyWorldBench,包含 100 个组装任务,涵盖 80 个物体,横跨家具、工业组装和断裂重组三个领域。对八个智能体系统的评估揭示了它们能力上的显著差异。最强的系统实现了 80.9% 的零件准确率,但完整组装成功率仅为 59.4%。被评估的开源系统在执行可靠性和组装准确率上都大幅落后于更强的闭源同行。对视觉参考、交互轨迹和失败案例的分析表明,智能体在修正组装的同时仍残留定位误差。AssemblyWorld 为评估交互式组装智能体的能力和刻画近似结构恢复与精确重建之间的差距提供了一个通用平台。
原文摘要
The task of 3D assembly requires translating an understanding of parts and their relationships into precise spatial arrangements. Can pretrained general-purpose agents assemble objects through visual interaction without additional assembly-specific fine-tuning? To investigate this question, we introduce AssemblyWorld, an interactive 3D environment in which agents inspect rendered views and manipulate supplied rigid parts, guided by images or assembly manuals when available. Agents perceive part geometry through 2D views rather than direct access to mesh vertices or faces, while their resulting assemblies are evaluated geometrically. Building on this environment, we construct AssemblyWorldBench, comprising 100 assembly tasks across 80 objects spanning furniture, industrial assembly, and fra...
*自动采集于 2026-10-02*
#论文 #arXiv #CV #小凯