论文概要
研究领域: CV
作者: Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu
发布时间: 2026-09-22
arXiv: 2609.26793
中文摘要
组合式 3D 场景重建近来沿两个方向探索:提供空间关系语义理解的智能体推理,但缺乏与输入图像的精确对齐;从输入图像预测稠密点图的视觉几何基础模型,但重建质量有限。从单张单目图像恢复"物体关系准确 + 重建高保真"的完整 3D 场景依然困难。本文提出 HARMONY——分层思维链框架,兼用智能体推理与视觉几何。给定室内场景图像,从空 3D 平面图出发:先对照参考图像标定相机,建立语义落地的空间坐标系;再用智能体 VLM 推理恢复 3D 房间布局与初始摆放顺序;随后按层级摆放物体——从墙面挂件、独立家具,到家具顶部从属装饰。家具采用深度优先遍历,每次摆放以已确立结构为条件,并配反思反馈回路避免误差累积。每次 VLM 摆放后,再用点云估计做几何精调,使渲染图像更贴合输入。HARMONY 生成语义一致、感知上与参考图像对齐的 3D 场景,将单图组合重建扩展到复杂室内场景。合成与真实图像实验显示其优于所评测基线;与 GPT-6 Astra 的定性比较表明其物体布置更忠实、场景细节保留更好。
原文摘要
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality remains challenging. In this paper, we present HARMONY, a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Given an image of an indoor scene, starting from an empty 3D floorplan, HARMONY first calibrates the camera against the reference image to establish a semantically-grounded spatial frame, then uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. It then places the objects in a hierarchical order, from wall-mounted elements, free-standing furniture, to dependent decorations on top of furniture. We also use depth-first traversal for furniture so each placement conditions on previously resolved structure and a reflective feedback loop to avoid error accumulation. After each object placement by VLM, we use the point cloud estimations to perform geometry-based refinement so that the rendered image aligns better with the input. HARMONY can produce 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.
自动采集于 2026-09-24
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。