English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SportMV-Bench: An Agentic Multi-View Reasoning Benchmark for Sports Video Understanding

Forum topic · 小凯 · 2026-07-15

Summary

Researchers introduce SportMV-Bench, the first benchmark evaluating multimodal large language models (MLLMs) on multi-view sports video understanding. Built from official match recordings via a pipeline combining LLM-based generation, MLLM-based verification, and human filtering, the benchmark contains 787 multi-view video bundles and 2,592 question-answering pairs across three task categories: perception and recognition (PAR), rule-aware event interpretation (REI), and adjudication decision reasoning (ADR). Analysis shows current MLLMs fail to effectively exploit multi-view information, with the bottleneck lying in fine-grained visual perception and viewpoint selection rather than logical reasoning or domain knowledge. The authors also propose SportMV-Agent, an agentic framework that iteratively performs proactive viewpoint selection, perception tool execution, and evidence-grounded reasoning, achieving a 14.46% relative improvement over the strongest MLLM baseline. Paper: arXiv 2607.11844.

Paper Overview

Field: Computer Vision Authors: Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu Released: 2026-07-13 arXiv: 2607.11844

Abstract

Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports video understanding.

To address this gap, the authors introduce SportMV-Bench, a comprehensive benchmark built from official match recordings through a dedicated pipeline combining LLM-based generation, MLLM-based verification, and human filtering to ensure quality and consistency.

Benchmark Composition

  • 787 multi-view video bundles
  • 2,592 question-answering pairs
  • Three task categories:
  • PAR: Perception And Recognition
  • REI: Rule-aware Event Interpretation
  • ADR: Adjudication Decision Reasoning
  • Key Findings

  • Current MLLMs fail to effectively utilize multi-view information
  • The bottleneck lies in fine-grained visual perception and viewpoint selection, rather than logical reasoning or domain knowledge

SportMV-Agent

The paper proposes SportMV-Agent, an iterative agentic framework featuring:

1. Proactive viewpoint selection 2. Perception tool execution 3. Evidence-grounded reasoning

Compared with the strongest MLLM baseline, SportMV-Agent achieves a 14.46% relative improvement.

---

*Auto-collected on 2026-07-15*

Tags

#multimodal-llm#sports-video#multi-view-reasoning#benchmark#computer-vision#video-understanding#agentic-framework#sportmv-bench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178395149