Overview
Field: Computer Vision Authors: Ipek Oztas, Duygu Ceylan, Aybars Bugra Aksoy Published: 2026-06-27 arXiv: 2606.27332
Abstract
Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. The authors address this problem by introducing a geometry-aware object motion method that operates directly on the positional representations of diffusion transformers.
Key Ideas
- Problem: Relocating objects in single images demands consistent handling of occlusions, previously hidden regions, shadows, and reflections—a task existing editing methods struggle with.
- Approach: RoPEMover performs object motion directly in the positional embedding space of diffusion transformers, making the relocation geometry-aware (leveraging depth cues).
- Benefit: Scene-level consistency—occlusion relationships, lighting, and reflections—is better preserved than with prior pixel- or latent-space editing methods.
*Auto-collected on 2026-06-27*