English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HARMONY: Hierarchical Agentic Reasoning for Monocular 3D Indoor Scene Reconstruction

Forum topic · 小凯 · 2026-09-24

Summary

HARMONY is a hierarchical chain-of-thought framework that combines agentic VLM reasoning with visual geometry foundation models to reconstruct complete 3D indoor scenes from a single monocular image. Existing approaches either offer semantic understanding of spatial relationships without precise image alignment (agentic reasoning) or predict dense point maps with limited reconstruction quality (visual geometry models). HARMONY addresses this gap by starting from an empty 3D floorplan: it calibrates the camera against the reference image to establish a semantically grounded spatial frame, recovers the 3D room layout and placement order via agentic VLM reasoning, and places objects hierarchically—from wall-mounted elements to free-standing furniture to dependent decorations. Depth-first traversal conditions each placement on resolved structure, while a reflective feedback loop prevents error accumulation. Point-cloud-based geometric refinement after each placement keeps rendered views aligned with the input. Experiments on synthetic and real-world images show HARMONY outperforms evaluated baselines, with qualitative comparisons against GPT-6 Astra indicating more faithful object arrangements and better scene detail preservation.

Overview

Field: Computer Vision (CV) Authors: Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu Published: 2026-09-22 arXiv: 2609.26793

Introduction

Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning, which provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models, which predict dense point maps from input images but offer limited reconstruction quality. Recovering a complete 3D scene from a single monocular image—with accurate inter-object relationships and high-fidelity reconstruction—remains challenging.

The HARMONY Framework

HARMONY is a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation models. Given an image of an indoor scene, starting from an empty 3D floorplan:

1. Camera calibration: HARMONY first calibrates the camera against the reference image to establish a semantically grounded spatial frame. 2. Layout reasoning: It uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. 3. Hierarchical placement: Objects are placed in hierarchical order, from wall-mounted elements, to free-standing furniture, to dependent decorations on top of furniture. 4. Structured traversal: Depth-first traversal for furniture ensures each placement conditions on previously resolved structure, with a reflective feedback loop to avoid error accumulation. 5. Geometric refinement: After each object placement by the VLM, point cloud estimations perform geometry-based refinement so that the rendered image aligns better with the input.

Results

HARMONY produces 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.

--- *Auto-collected on 2026-09-24*

Tags

#3d-reconstruction#computer-vision#monocular-image#vlm#agentic-reasoning#scene-generation#arxiv#chain-of-thought

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635139