English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Referring Expressions: Scenario Comprehension Visual Grounding (RSC Benchmark & ScenGround)

Forum topic · 小凯 · 2026-04-04

Summary

This paper introduces Referring Scenario Comprehension (RSC), a new benchmark for scenario-based visual grounding where targets must be inferred from object roles, user intentions, and relational context rather than matched to literal referring expressions. RSC queries are paragraph-length texts with deliberate distractor references, and each instance carries interpretable difficulty tags (uniqueness, clutter, size, overlap, position) for fine-grained analysis. The dataset includes about 31k training samples, 4k in-domain test samples, and a 3k out-of-distribution split with unseen object categories. The authors also present ScenGround, a curriculum reasoning approach combining supervised warm-up with difficulty-aware reinforcement learning. Experiments show scenario-based queries expose systematic failures in current models that standard benchmarks miss, and curriculum training improves performance on challenging slices while transferring to standard benchmarks. arXiv: 2504.01259.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Ruozhen He, Nisarg A. Shah, Qihua Dong
  • Released: 2025-04-01
  • arXiv: 2504.01259
  • Abstract (translated from the Chinese summary)

    Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. This work explores a complementary and more challenging setting of scenario-based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming.

    The authors introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting. Queries are paragraph-length texts describing object roles, user goals, and contextual cues, including deliberate references to distractor objects that often require deep understanding to resolve. Each instance is annotated with interpretable difficulty tags — uniqueness, clutter, size, overlap, and position — which expose distinct failure modes and support fine-grained analysis.

    RSC contains:

  • ~31k training samples
  • 4k in-domain test samples
  • A 3k out-of-distribution split with unseen object categories
The authors further propose ScenGround, a curriculum reasoning method serving as a reference point for this setting, combining supervised warm-up with difficulty-aware reinforcement learning.

Experiments show that scenario-based queries expose systematic failures in current models that standard benchmarks do not reveal, and that curriculum training improves performance on challenging slices while transferring to standard benchmarks.

Original Abstract (excerpt)

> Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent named category. We explore a complementary and more challenging setting of scenario-based visual grounding, where the target must be inferred from roles, intentions, and relational context rather than explicit naming. We introduce Referring Scenario Comprehension (RSC), a benchmark designed for this setting...

---

*Auto-collected on 2026-04-04*

Tags

#computer-vision#visual-grounding#benchmark#reinforcement-learning#multimodal#arxiv#scenground

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169525