English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

VideoDetective: Clue Hunting via Extrinsic Query and Intrinsic Affinity for Long-Video QA

Forum topic · 小凯 · 2026-03-25

Summary

VideoDetective is a framework for long-video question answering with multimodal large language models (MLLMs). Because MLLMs have limited context windows, answering questions over long videos requires identifying sparse query-relevant segments. Existing approaches localize clues based only on the query, ignoring the video's intrinsic structure and varying relevance across segments. VideoDetective divides a video into segments and represents them as a visual-temporal affinity graph built from visual similarity and temporal proximity. It then runs a Hypothesis-Verification-Refinement loop: relevance scores of observed segments are estimated with respect to the query and propagated through the graph to unobserved segments, producing a global relevance distribution that guides selection of the most critical segments for sparse observation and final answering. Experiments show consistent significant improvements across mainstream MLLMs on representative benchmarks, with up to 7.5% accuracy gain on VideoMME-long. Authors: Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu. arXiv: 2603.22285.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu
  • Published: 2026-03-23
  • arXiv: 2603.22285
  • Abstract

    Long video understanding remains challenging for multimodal large language models (MLLMs) due to limited context windows, which necessitate identifying sparse query-relevant video segments. However, existing methods predominantly localize clues based solely on the query, overlooking the video's intrinsic structure and varying relevance across segments.

    To address this, the authors propose VideoDetective, a framework that integrates query-to-segment relevance and inter-segment affinity for effective clue hunting in long-video question answering.

    Method

    1. Segmentation & graph construction: Divide the video into segments and represent them as a visual-temporal affinity graph built from visual similarity and temporal proximity. 2. Hypothesis-Verification-Refinement loop: Estimate relevance scores of observed segments with respect to the query, then propagate these scores to unobserved segments via the graph. 3. Sparse observation: The resulting global relevance distribution guides the selection of the most critical segments for the final answer.

    Results

  • Consistent, significant improvements across various mainstream MLLMs on representative benchmarks.
  • Accuracy gains of up to 7.5% on VideoMME-long.
  • Links

  • arXiv: https://arxiv.org/abs/2603.22285
---

*Auto-collected on 2026-03-25*

Tags

#long-video-understanding#multimodal-llm#video-question-answering#computer-vision#arxiv#graph-based-reasoning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169025