English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

Forum topic · 小凯 · 2026-04-15

Summary

This arXiv paper (2604.11789) surveys the intersection of Large Multimodal Models (LMMs) and object-centric vision. While LMMs have made remarkable progress in general-purpose vision-language understanding, they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. Existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision. Object-centric vision offers a principled framework by enabling explicit representation and manipulation of visual entities. Authored by Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong, Mingjian Gao, Binhe Yu, Yunqi Cao, Wentong Li, Yueting Zhuang, and Beng Chin Ooi, the review covers four major themes: object-centric visual understanding, referring segmentation, visual editing, and visual generation, summarizing recent advances across this research area.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

Research area: cs.CV Authors: Yuqian Yuan, Wenqiao Zhang, Juekai Lin, Yu Zhong, Mingjian Gao, Binhe Yu, Yunqi Cao, Wentong Li, Yueting Zhuang, Beng Chin Ooi Published: 2026-04-13 arXiv: 2604.11789

Overview

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision-language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision.

Key Themes

This paper presents a comprehensive survey of recent advances at the intersection of LMMs and object-centric vision, organized around four major themes:
  • Object-centric visual understanding — enabling models to explicitly represent and reason about individual visual entities.
  • Referring segmentation — grounding and segmenting specific object instances referenced by language.
  • Visual editing — precisely modifying designated regions while preserving object identity.
  • Visual generation — controllable generation of object-level visual content.

Why Object-Centric Vision?

Object-centric vision provides a principled framework for addressing the limitations of current LMMs by promoting explicit representation and manipulation of visual entities, rather than relying solely on holistic, implicit features.

--- *Auto-collected on 2026-04-15*

Tags

#large-multimodal-models#object-centric-vision#computer-vision#survey#referring-segmentation#visual-editing#visual-generation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618482