LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
Research area: cs.CV Authors: Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong, Mingjian Gao, Binhe Yu, Yunqi Cao, Wentong Li, Yueting Zhuang, Beng Chin Ooi Published: 2026-04-13 arXiv: 2604.11789
Overview
Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision-language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision.Key Themes
This paper presents a comprehensive survey of recent advances at the intersection of LMMs and object-centric vision, organized around four major themes:- Object-centric visual understanding — enabling models to explicitly represent and reason about individual visual entities.
- Referring segmentation — grounding and segmenting specific object instances referenced by language.
- Visual editing — precisely modifying designated regions while preserving object identity.
- Visual generation — controllable generation of object-level visual content.
Why Object-Centric Vision?
Object-centric vision provides a principled framework for addressing the limitations of current LMMs by promoting explicit representation and manipulation of visual entities, rather than relying solely on holistic, implicit features.--- *Auto-collected on 2026-04-15*