English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Searching Videos as Trees: VideoTreeSearch for Self-Correcting Grounded Long-Video QA

Forum topic · 小凯 · 2026-07-21

Summary

Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while also localizing the short evidence interval supporting the answer. Existing agentic methods use only a crop_video(start, end) action, enabling coarse-to-fine narrowing but lacking explicit backtracking, so agents often converge prematurely. VideoTreeSearch (VTS) reformulates grounded LVQA as iterative self-correcting search over an adaptive temporal tree built from visual scene boundaries, where each node is a semantically coherent segment. An agent navigates via four discrete operations: zoom_in, zoom_out, shift, and answer, making recovery an explicit, learnable primitive. Training combines a trajectory synthesis pipeline (including deliberate detours and recovery), supervised fine-tuning, and reinforcement learning with grounding and answer-accuracy rewards. On CG-Bench, Haystack-LVBench, and Haystack-Ego4D, VTS exceeds prior agentic methods by +12.5 mIoU and +7.4 T-F1, and transfers to general long-video QA (Video-MME, MLVU, LVBench) with gains up to +7.1 accuracy points.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

  • Research area: cs.CV
  • Authors: Ce Zhang, Ziyang Wang, Yulu Pan
  • Published: 2026-07-21
  • arXiv: 2507.15489
  • Overview

    Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine narrowing but provides no primitive for fine-to-coarse backtracking. As a result, these agents typically converge prematurely and cannot recover from an early mistake.

    VideoTreeSearch (VTS)

    VTS casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree:

  • A non-uniform tree is constructed from visual scene boundaries so that each node corresponds to a semantically coherent segment.
  • The agent navigates the tree through four discrete operations: zoom_in, zoom_out, shift, and answer, exposing backtracking and recovery as explicit, learnable primitives rather than implicit behaviors.
  • A trajectory synthesis pipeline produces multi-step paths through the tree, including deliberate detours into incorrect branches followed by recovery.
  • Training uses supervised fine-tuning on these trajectories, followed by reinforcement learning with grounding and answer-accuracy rewards.
  • Results

  • On three Grounded LVQA benchmarks (CG-Bench, Haystack-LVBench, Haystack-Ego4D), VTS outperforms the strongest prior agentic methods by +12.5 mIoU on CG-Bench and +7.4 T-F1 on Haystack-Ego4D.
  • The learned policy transfers to general long-video QA, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to +7.1 accuracy points.
  • Ablations confirm that self-correcting hierarchical search is the central mechanism behind these gains.

Tags

#video-question-answering#long-video-understanding#agentic-search#video-grounding#reinforcement-learning#vision-language-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446968