Paper Overview
- Field: NLP
- Authors: Shaoxuan Li, Zhixuan Zhao, Hanze Deng
- Published: 2025-03-30
- arXiv: 2503.23716
Abstract
We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is sufficient: answering each question requires multiple temporally separated pieces of visual evidence and compositional constraints under conjunctive and sequential logic, spanning perceptual subtasks such as objects, attributes, relations, locations, actions, and events, and requiring skills including semantic recognition, visual correspondence, temporal reasoning, and spatial reasoning.
The benchmark contains 1,114 highly complex questions over 279 videos from diverse domains, including city walking tours, indoor house tours, video games, and extreme outdoor sports, 100% manually annotated.
Human studies show that PerceptionComp requires substantial test-time thinking and repeated perception steps: participants spent longer than on previous benchmarks, and accuracy dropped to near-random (18.97%) when re-watching was prohibited.
State-of-the-art multimodal large language models also perform far worse on PerceptionComp than on existing benchmarks: the best model in our evaluation, Gemini-3-Flash, achieved only 45.96% accuracy in a five-choice setting, while open-source models remained below 40%.
These results indicate that perception-centric long-horizon video reasoning remains a major bottleneck, and the authors hope PerceptionComp will help drive progress in perceptual reasoning.
---
*Auto-collected on 2026-03-31.*