Paper Overview
Field: Computer Vision Authors: Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu Released: 2026-07-13 arXiv: 2607.11844
Abstract
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports video understanding.
To address this gap, the authors introduce SportMV-Bench, a comprehensive benchmark built from official match recordings through a dedicated pipeline combining LLM-based generation, MLLM-based verification, and human filtering to ensure quality and consistency.
Benchmark Composition
- 787 multi-view video bundles
- 2,592 question-answering pairs
- Three task categories:
- PAR: Perception And Recognition
- REI: Rule-aware Event Interpretation
- ADR: Adjudication Decision Reasoning
- Current MLLMs fail to effectively utilize multi-view information
- The bottleneck lies in fine-grained visual perception and viewpoint selection, rather than logical reasoning or domain knowledge
Key Findings
SportMV-Agent
The paper proposes SportMV-Agent, an iterative agentic framework featuring:
1. Proactive viewpoint selection 2. Perception tool execution 3. Evidence-grounded reasoning
Compared with the strongest MLLM baseline, SportMV-Agent achieves a 14.46% relative improvement.
---
*Auto-collected on 2026-07-15*