Overview
Field: Computer Vision (CV) Authors: Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu Published: 2026-09-22 arXiv: 2609.26793
Introduction
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning, which provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models, which predict dense point maps from input images but offer limited reconstruction quality. Recovering a complete 3D scene from a single monocular image—with accurate inter-object relationships and high-fidelity reconstruction—remains challenging.
The HARMONY Framework
HARMONY is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation models. Given an image of an indoor scene, starting from an empty 3D floorplan:
1. Camera calibration: HARMONY first calibrates the camera against the reference image to establish a semantically grounded spatial frame. 2. Layout reasoning: It uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. 3. Hierarchical placement: Objects are placed in hierarchical order, from wall-mounted elements, to free-standing furniture, to dependent decorations on top of furniture. 4. Structured traversal: Depth-first traversal for furniture ensures each placement conditions on previously resolved structure, with a reflective feedback loop to avoid error accumulation. 5. Geometric refinement: After each object placement by the VLM, point cloud estimations perform geometry-based refinement so that the rendered image aligns better with the input.
Results
HARMONY produces 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.
--- *Auto-collected on 2026-09-24*