Paper Overview
- Field: Computer Vision (CV)
- Authors: Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu
- Published: 2026-03-23
- arXiv: 2603.22285
- Consistent, significant improvements across various mainstream MLLMs on representative benchmarks.
- Accuracy gains of up to 7.5% on VideoMME-long.
- arXiv: https://arxiv.org/abs/2603.22285
Abstract
Long video understanding remains challenging for multimodal large language models (MLLMs) due to limited context windows, which necessitate identifying sparse query-relevant video segments. However, existing methods predominantly localize clues based solely on the query, overlooking the video's intrinsic structure and varying relevance across segments.
To address this, the authors propose VideoDetective, a framework that integrates query-to-segment relevance and inter-segment affinity for effective clue hunting in long-video question answering.
Method
1. Segmentation & graph construction: Divide the video into segments and represent them as a visual-temporal affinity graph built from visual similarity and temporal proximity. 2. Hypothesis-Verification-Refinement loop: Estimate relevance scores of observed segments with respect to the query, then propagate these scores to unobserved segments via the graph. 3. Sparse observation: The resulting global relevance distribution guides the selection of the most critical segments for the final answer.
Results
Links
*Auto-collected on 2026-03-25*