English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

3D-Layout-R1: Structured Reasoning for Language-Guided Spatial Layout Editing

Forum topic · 小凯 · 2026-03-25

Summary

3D-Layout-R1 (arXiv:2603.22279) introduces a structured reasoning framework that performs text-conditioned spatial layout editing via scene-graph reasoning. While large language models (LLMs) and vision-language models (VLMs) show impressive reasoning abilities, they struggle with spatial understanding and layout consistency in fine-grained visual editing. Given an input scene graph and a natural-language instruction, the model reasons over the graph to produce an updated scene graph that satisfies the text condition while maintaining spatial coherence. Explicitly guiding reasoning with structured relational representations improves interpretability and control over spatial relationships. Evaluated on a new text-guided layout editing benchmark covering ordering, spatial alignment, and room editing tasks, the method outperforms chain-of-thought SFT (CoT-SFT) and vanilla GRPO baselines, achieving an average 15% IoU improvement and 25% reduction in center distance error. Authors include Haoyu Zhen, Xiaolong Li, Yilin Zhao, Han Zhang, Sifei Liu, Kaichun Mo, Chuang Gan, and Subhashree Radhakrishnan.

Paper Overview

Field: Computer Vision (CV) Authors: Haoyu Zhen, Xiaolong Li, Yilin Zhao, Han Zhang, Sifei Liu, Kaichun Mo, Chuang Gan, Subhashree Radhakrishnan Published: 2026-03-23 arXiv: 2603.22279

Abstract

Large Language Models (LLMs) and Vision Language Models (VLMs) have shown impressive reasoning abilities, yet they struggle with spatial understanding and layout consistency when performing fine-grained visual editing. This paper introduces a Structured Reasoning framework that performs text-conditioned spatial layout editing via scene-graph reasoning.

Given an input scene graph and a natural-language instruction, the model reasons over the graph to generate an updated scene graph that satisfies the text condition while maintaining spatial coherence. By explicitly guiding the reasoning process through structured relational representations, the approach improves both interpretability and control over spatial relationships.

Evaluation

The method is evaluated on a new text-guided layout editing benchmark covering:

  • Ordering tasks
  • Spatial alignment tasks
  • Room editing tasks
  • Compared with chain-of-thought supervised fine-tuning (CoT-SFT) and vanilla GRPO baselines, the proposed training paradigm achieves:

  • An average 15% improvement in IoU
  • A 25% reduction in center distance error
  • Key Takeaways

  • Scene-graph-based structured reasoning provides an explicit, interpretable intermediate representation for spatial editing.
  • Conditioning reasoning on structured relational representations improves spatial consistency over free-form chain-of-thought approaches.
  • The introduced benchmark offers a standardized testbed for text-guided layout editing.
---

*Auto-collected on 2026-03-25*

Tags

#3d-layout-r1#scene-graph#spatial-reasoning#layout-editing#vision-language-models#reinforcement-learning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169030