English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Thinking in Boxes: Precise 3D Image Editing in Real Images via Box Specifications

Forum topic · 小凯 · 2026-06-20

Summary

Thinking in Boxes (arXiv:2506.16804) is a computer vision paper proposing a 3D box-based interface for editing real images. Instead of ambiguous text or 2D conditioning, users specify an edit by providing input and output 3D boxes, turning image editing into a well-posed geometry problem. Color-coded box faces convey 3D orientation, enabling precise control over translation, rotation, scaling, and viewpoint changes while preserving scene and object identity and recovering previously unseen object regions. The method grounds transformations in scene appearance using a depth-aligned planar floor as a global reference frame with depth-aware coloring. An image generator conditioned on this structure is trained in two stages—first on synthetic multi-object scenes, then on a small set of real Objectron videos—achieving generalization to complex in-the-wild real images. It outperforms recent state-of-the-art methods on large 3D edits of real photos.

Overview

Field: Computer Vision Authors: Pradhaan S Bhat, Naveen Chandra R, Rishubh Parihar Published: 2025-06-20 arXiv: 2506.16804

Text- and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing, particularly under large object motions and camera changes. Prior work has used 3D primitives such as boxes only as loose conditioning signals indicating approximate object location, rather than specifying the transformation.

This paper instead uses 3D boxes as structured specifications: the user provides the input and output boxes of the edit, casting editing as a well-posed geometry problem. This "thinking in boxes" interface, where each box face is color-coded to convey 3D orientation, gives precise control over translation, rotation, scaling, and viewpoint changes in real images while preserving scene and object identity, and recovering previously unseen object regions.

To ground transformations in scene appearance, the authors introduce a depth-aligned planar floor as a global reference frame, colored with depth-aware cues. An image generator conditioned on this structure produces consistent results under large transformations. The system is trained in two stages: first on synthetic multi-object scenes, then on a small number of real Objectron videos, ultimately generalizing to complex in-the-wild real images. The method operates directly on real photographs and substantially outperforms recent state-of-the-art methods on large 3D edits.

Original Abstract

Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as boxes, but only as loose conditioning signals indicating approximate object location rather than specifying the transformation. We instead use 3D boxes as structured specifications: the user provides the input and output boxes of the edit, casting editing as a well-posed geometry problem. This 'thinking in boxes' interface, where each box face is color-coded to convey 3D orientation, gives precise control over translation, rotation, scaling, and viewpoint changes in real images while preserving scene and object identity, and recovering previously unseen object regions.

--- *Auto-collected on 2026-06-20*

Tags

#computer-vision#image-editing#3d-boxes#generative-models#arxiv-paper#3d-editing#viewpoint-control

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981551