English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Seek to Segment: Active Perception for Panoramic Referring Segmentation (PanoSeeker)

Forum topic · 小凯 · 2026-07-06

Summary

This paper introduces Active Panoramic Referring Segmentation (APRS), a new embodied AI task in which an agent must actively adjust its viewing direction (Δθ, Δφ) to explore a continuous 360° environment and locate the object specified by a user instruction for segmentation. Unlike existing referring segmentation models that passively process static images from fixed viewpoints, APRS requires efficient, goal-directed visual search. The authors propose PanoSeeker, a memory-augmented agent that combines a Vision-Language Model (VLM) with EgoSphere, an explicit spatial visual memory. EgoSphere progressively integrates sequential local observations into a unified 360° representation, enabling the agent to plan efficient, non-redundant search trajectories instead of heuristic scanning. Once the target is found, the agent performs active viewpoint alignment and outputs the segmentation mask. The work also contributes an expert-annotated dataset of search trajectories with memory timelines for supervised fine-tuning, followed by reinforcement learning post-training to explicitly optimize exploration efficiency. Experiments on a newly established APRS benchmark show that PanoSeeker achieves superior search efficiency and segmentation accuracy, significantly outperforming adapted state-of-the-art baselines. Paper: arXiv:2607.02497.

Overview

Field: Computer Vision Authors: Song Tang, Shuming Hu, Xincheng Shuai, Henghui Ding, Yu-Gang Jiang arXiv: 2607.02497 Category: cs.CV

Key Contributions

  • New Task — Active Panoramic Referring Segmentation (APRS): Existing referring segmentation models passively process static images from fixed perspectives, limiting their applicability in Embodied AI, where agents must perform active perception in continuous 360° environments. In APRS, an agent adjusts its viewing direction (Δθ, Δφ) to explore the 360° environment and seek the object specified by a user instruction for segmentation.
  • PanoSeeker, a memory-augmented agent: Rather than relying on heuristic scanning, PanoSeeker integrates a Vision-Language Model (VLM) with EgoSphere, an explicit spatial visual memory. By progressively integrating sequential local observations into a unified 360° representation, EgoSphere enables the agent to plan efficient and non-redundant search trajectories. Once the target is found, the agent performs active viewpoint alignment and outputs the segmentation mask.
  • Training data and post-training: The authors curate an expert-annotated dataset of search trajectories with memory timelines for supervised fine-tuning, followed by reinforcement learning post-training to explicitly optimize exploration efficiency.

Results

Extensive experiments on the newly established APRS benchmark demonstrate that PanoSeeker achieves excellent performance in both search efficiency and segmentation accuracy, significantly outperforming adapted state-of-the-art baseline methods.

---

*Auto-collected on 2026-07-06*

Tags

#computer-vision#embodied-ai#referring-segmentation#active-perception#360-degree#vision-language-model#reinforcement-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178209071