English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RoPEMover: Depth-Aware Object Relocation via Positional Embeddings in Diffusion Transformers

Forum topic · 小凯 · 2026-06-27

Summary

RoPEMover is a computer vision method for moving objects within a single image while maintaining geometry-consistent scene layout. Object relocation requires handling occlusions, revealing previously unseen regions, and preserving coherent shadows and reflections—tasks where existing approaches often fail to maintain scene-level consistency. Proposed by Ipek Oztas, Duygu Ceylan, and Aybars Bugra Aksoy, the method introduces a geometry-aware object motion approach that operates directly on the positional representations (rotary positional embeddings) of diffusion transformers. Instead of manipulating pixels or latent features alone, RoPEMover leverages depth information and positional embeddings to rearrange objects in a geometrically consistent way, preserving lighting effects and visibility relationships in the edited scene. The paper is available on arXiv as 2606.27332.

Overview

Field: Computer Vision Authors: Ipek Oztas, Duygu Ceylan, Aybars Bugra Aksoy Published: 2026-06-27 arXiv: 2606.27332

Abstract

Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. The authors address this problem by introducing a geometry-aware object motion method that operates directly on the positional representations of diffusion transformers.

Key Ideas

  • Problem: Relocating objects in single images demands consistent handling of occlusions, previously hidden regions, shadows, and reflections—a task existing editing methods struggle with.
  • Approach: RoPEMover performs object motion directly in the positional embedding space of diffusion transformers, making the relocation geometry-aware (leveraging depth cues).
  • Benefit: Scene-level consistency—occlusion relationships, lighting, and reflections—is better preserved than with prior pixel- or latent-space editing methods.
---

*Auto-collected on 2026-06-27*

Tags

#computer-vision#diffusion-transformers#image-editing#object-relocation#positional-embeddings#rope#depth-awareness#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208208