Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
- Research area: cs.CV
- Authors: Ce Zhang, Ziyang Wang, Yulu Pan
- Published: 2026-07-21
- arXiv: 2507.15489
- A non-uniform tree is constructed from visual scene boundaries so that each node corresponds to a semantically coherent segment.
- The agent navigates the tree through four discrete operations:
zoom_in,zoom_out,shift, andanswer, exposing backtracking and recovery as explicit, learnable primitives rather than implicit behaviors. - A trajectory synthesis pipeline produces multi-step paths through the tree, including deliberate detours into incorrect branches followed by recovery.
- Training uses supervised fine-tuning on these trajectories, followed by reinforcement learning with grounding and answer-accuracy rewards.
- On three Grounded LVQA benchmarks (CG-Bench, Haystack-LVBench, Haystack-Ego4D), VTS outperforms the strongest prior agentic methods by +12.5 mIoU on CG-Bench and +7.4 T-F1 on Haystack-Ego4D.
- The learned policy transfers to general long-video QA, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to +7.1 accuracy points.
- Ablations confirm that self-correcting hierarchical search is the central mechanism behind these gains.
Overview
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine narrowing but provides no primitive for fine-to-coarse backtracking. As a result, these agents typically converge prematurely and cannot recover from an early mistake.
VideoTreeSearch (VTS)
VTS casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree: