Summary
A 2026 arXiv paper (2607.09654) by Shravan Murlidaran and Miguel P. Eckstein evaluates the evolution of vision-language models (VLMs) in scene description over 2017–2025. While most prior evaluations use simple MS-COCO scenes, the authors introduce the Complex Social Behavior (CSB) dataset: 100 images depicting intricate social interactions and behaviors. Analyzing four pre-MLLM and five modern MLLM models, they find that accuracy gains on CSB are even larger than on MS-COCO. Pre-MLLM models perform far below humans, whereas modern MLLMs reach parity with top human performers. MLLMs nearly eliminate all error types except occasional spatial dependency errors—relying on image regions that differ from those humans use. This work offers a benchmark and error taxonomy for assessing visual-cognitive capabilities of multimodal models on complex social scenes.
Paper Overview
Field: CV/AI
Authors: Shravan Murlidaran, Miguel P. Eckstein
Published: 2026-07-10
arXiv: 2607.09654
Abstract (English translation)
Vision-language models (VLMs) have made remarkable progress in visual reasoning over the past decade, but most evaluations rely on simple scenes (e.g., MS-COCO). This paper introduces the Complex Social Behavior (CSB) dataset, consisting of 100 images depicting complex social interactions and behaviors. The authors analyze the progress of scene description by VLMs from 2017 to 2025, covering 4 pre-MLLM and 5 modern MLLM models.
Key findings
- The CSB dataset reveals even more pronounced accuracy improvements than MS-COCO.
- Pre-MLLM models perform far below human level on these complex social scenes.
- Modern MLLMs achieve accuracy on par with top human performers.
- MLLMs have almost eliminated all error types, except occasionally relying on image regions that differ from those humans depend on (spatial dependency errors).
---
*Auto-collected on 2026-07-14.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178395111