English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Thinking in Boxes: 3D Editing in Real Images Made Easy

Forum topic · 小凯 · 2026-06-22

Summary

Researchers Pradhaan S Bhat, Naveen Chandra R, and Rishubh Parihar introduce 'Thinking in Boxes,' a 3D image editing framework (arXiv 2506.17584) that treats 3D bounding boxes as a structured canonical interface rather than a loose conditioning signal. Users specify an input box and an output box for an object, converting edits like translation, rotation, scaling, and viewpoint changes into a well-defined geometric problem. Each box face is color-coded for 3D orientation, giving precise control over large transformations in real photographs while preserving scene and object identity and recovering previously unseen regions. A depth-aligned plane floor, shaded with depth-aware cues, serves as a global reference frame to anchor transformations to scene appearance. The model is trained in two stages: pretraining on synthetic multi-object scenes, then fine-tuning on a small subset of real-world videos from the Objectron dataset. The approach generalizes to complex in-the-wild photos and reportedly outperforms recent state-of-the-art methods on large 3D editing tasks.

Paper Overview

Field: cs.CV Authors: Pradhaan S Bhat, Naveen Chandra R, Rishubh Parihar Published: 2026-06-21 arXiv: 2506.17584

Abstract (Translation)

Text- and 2D-based conditioning interfaces offer weak and ambiguous control over spatial transformations in image editing—especially under large object motions or changes in camera viewpoint. Prior work has introduced 3D primitives (such as cuboids) but only as loose conditioning signals that roughly indicate object location, rather than precisely specifying the transformation itself.

This work takes the opposite approach: it uses 3D boxes as a structured canonical. The user provides an input box (before the edit) and an output box (after the edit), turning the entire editing process into a well-defined geometric problem. In this "Thinking in Boxes" interface, each face is color-coded with a distinct color to encode 3D orientation, giving users precise control over translation, rotation, scaling, and even viewpoint transformations in real images. The method also preserves scene and object identity and recovers previously unseen object regions.

To anchor the transformation firmly to scene appearance, the authors introduce a depth-aligned plane floor serving as a global reference frame, shaded with depth-aware cues. Guided by this structure, the image generator produces highly consistent results even under large transformations.

The system is trained in two stages: first pretrained on synthetic multi-object scenes, then fine-tuned on a small subset of real-world videos from the Objectron dataset, ultimately generalizing to complex, in-the-wild real photographs. The method operates directly on real photos and significantly outperforms recent state-of-the-art approaches on large 3D editing tasks.

Plain-Language Explanation

Imagine you want to move a table in a photo. Previous methods were like pointing directions through frosted glass—"probably around here." This method gives you two transparent 3D boxes: one marking "now" and one marking "where it should go." Each box face is painted a different color to indicate front/back/left/right/up/down, as clear as building with LEGO blocks. The floor also measures depth, like laying invisible coordinate paper across the room. This way the AI understands exactly how you want to "relocate" the object—and the result looks real, with no visual artifacts.

--- *Auto-collected on 2026-06-21*

Tags

#computer-vision#image-editing#3d-editing#generative-ai#deep-learning#arxiv#objectron#diffusion-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178207990