Paper Overview
Field: Computer Vision (CV) Authors: Jeffri Murrugarra-Llerena, Pranav Chitale, Zicheng Liu, Kai Ao, Yujin Ham, Guha Balakrishnan, Paola Cascante-Bonilla Published: 2026-03-23 arXiv: 2603.22249
Abstract
Social group detection, or the identification of humans involved in reciprocal interpersonal interactions (e.g., family members, friends, and customers and merchants), is a crucial component of social intelligence needed for agents transacting in the world. The few existing benchmarks for social group detection are limited by low scene diversity and reliance on third-person camera sources (e.g., surveillance footage). Consequently, these benchmarks generally lack real-world evaluation on how groups form and evolve in diverse cultural contexts and unconstrained settings.
To address this gap, the authors introduce EgoGroups, a first-person view dataset that captures social dynamics in cities around the world. EgoGroups spans 65 countries covering low, medium, and high-crowd settings under four weather/time-of-day conditions. The dataset provides dense human annotations of both individuals and social groups, along with rich geographic and scene metadata.
Using this dataset, the authors conduct extensive evaluations of state-of-the-art vision-language models (VLMs), large language models (LLMs), and supervised models on group detection tasks.
Key Contributions
- EgoGroups dataset: a large-scale egocentric benchmark for social group detection with global coverage
- Dense annotations: human and social group labels plus geographic and scene metadata
- Benchmarking: evaluation of VLMs, LLMs, and supervised baselines
*Source: arXiv:2603.22249*