English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

EgoGroups: A First-Person Benchmark for Detecting Social Groups of People in the Wild

Forum topic · 小凯 · 2026-03-25

Summary

EgoGroups is a new benchmark dataset for social group detection—identifying humans engaged in reciprocal interpersonal interactions such as families, friends, or customers and merchants. Existing benchmarks suffer from limited scene diversity and reliance on third-person cameras like surveillance footage, preventing realistic evaluation of how groups form and evolve across cultural contexts. EgoGroups addresses this with a first-person (egocentric) view dataset capturing social dynamics in cities worldwide, spanning 65 countries, low/medium/high crowd densities, and four weather/time-of-day conditions. The dataset includes dense human annotations of individuals and social groups, plus rich geographic and scene metadata. The authors benchmark state-of-the-art vision-language models, LLMs, and supervised models on group detection tasks. Paper by Murrugarra-Llerena et al., arXiv:2603.22249.

Paper Overview

Field: Computer Vision (CV) Authors: Jeffri Murrugarra-Llerena, Pranav Chitale, Zicheng Liu, Kai Ao, Yujin Ham, Guha Balakrishnan, Paola Cascante-Bonilla Published: 2026-03-23 arXiv: 2603.22249

Abstract

Social group detection, or the identification of humans involved in reciprocal interpersonal interactions (e.g., family members, friends, and customers and merchants), is a crucial component of social intelligence needed for agents transacting in the world. The few existing benchmarks for social group detection are limited by low scene diversity and reliance on third-person camera sources (e.g., surveillance footage). Consequently, these benchmarks generally lack real-world evaluation on how groups form and evolve in diverse cultural contexts and unconstrained settings.

To address this gap, the authors introduce EgoGroups, a first-person view dataset that captures social dynamics in cities around the world. EgoGroups spans 65 countries covering low, medium, and high-crowd settings under four weather/time-of-day conditions. The dataset provides dense human annotations of both individuals and social groups, along with rich geographic and scene metadata.

Using this dataset, the authors conduct extensive evaluations of state-of-the-art vision-language models (VLMs), large language models (LLMs), and supervised models on group detection tasks.

Key Contributions

  • EgoGroups dataset: a large-scale egocentric benchmark for social group detection with global coverage
  • Dense annotations: human and social group labels plus geographic and scene metadata
  • Benchmarking: evaluation of VLMs, LLMs, and supervised baselines
---

*Source: arXiv:2603.22249*

Tags

#computer-vision#social-group-detection#egocentric-vision#benchmark-dataset#vision-language-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169020