English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

Forum topic · 小凯 · 2026-03-31

Summary

PerceptionComp is a fully human-annotated benchmark for complex, long-horizon, perception-centric video reasoning, introduced in arXiv paper 2503.23716. Its design ensures that no single video moment suffices: each question requires combining multiple temporally separated visual evidence segments under conjunctive and sequential logic, spanning perceptual subtasks such as objects, attributes, relations, locations, actions, and events, and demanding semantic recognition, visual correspondence, temporal reasoning, and spatial reasoning. The benchmark contains 1,114 highly complex questions over 279 videos from diverse domains, including city walking tours, indoor house tours, video games, and extreme outdoor sports, all manually annotated. Human studies show the benchmark demands substantial test-time thinking and repeated perception: participants took longer than on prior benchmarks, and accuracy dropped to near-random (18.97%) when re-watching was disallowed. State-of-the-art multimodal LLMs struggle as well: the best model, Gemini-3-Flash, achieved only 45.96% accuracy in a five-choice setting, while open-source models remained below 40%. Results indicate perception-centric long-horizon video reasoning remains a major bottleneck for current models.

Paper Overview

  • Field: NLP
  • Authors: Shaoxuan Li, Zhixuan Zhao, Hanze Deng
  • Published: 2025-03-30
  • arXiv: 2503.23716

Abstract

We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is sufficient: answering each question requires multiple temporally separated pieces of visual evidence and compositional constraints under conjunctive and sequential logic, spanning perceptual subtasks such as objects, attributes, relations, locations, actions, and events, and requiring skills including semantic recognition, visual correspondence, temporal reasoning, and spatial reasoning.

The benchmark contains 1,114 highly complex questions over 279 videos from diverse domains, including city walking tours, indoor house tours, video games, and extreme outdoor sports, 100% manually annotated.

Human studies show that PerceptionComp requires substantial test-time thinking and repeated perception steps: participants spent longer than on previous benchmarks, and accuracy dropped to near-random (18.97%) when re-watching was prohibited.

State-of-the-art multimodal large language models also perform far worse on PerceptionComp than on existing benchmarks: the best model in our evaluation, Gemini-3-Flash, achieved only 45.96% accuracy in a five-choice setting, while open-source models remained below 40%.

These results indicate that perception-centric long-horizon video reasoning remains a major bottleneck, and the authors hope PerceptionComp will help drive progress in perceptual reasoning.

---

*Auto-collected on 2026-03-31.*

Tags

#video-benchmark#multimodal-llm#perception-reasoning#arxiv-paper#computer-vision#nlp#model-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169449