English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniAgent: Native Active Perception as Reasoning for Omni-Modal Long Video Understanding

Forum topic · 小凯 · 2026-06-19

Summary

OmniAgent is the first native omni-modal agent that formulates long video understanding as a POMDP-based iterative Observation-Thought-Action cycle, replacing the passive watch-it-all paradigm. Instead of uniformly processing all frames, OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, decoupling reasoning complexity from raw video duration. The framework introduces Agentic Supervised Fine-Tuning, which bootstraps native active perception via best-of-N trajectory synthesis with dual-stage quality control, and Agentic Reinforcement Learning with TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which uses turn-level entropy to steer credit assignment toward pivotal discovery turns. OmniAgent exhibits positive test-time scaling, improving as reasoning turns increase. Experiments across ten benchmarks including VideoMME and LVBench show state-of-the-art open-source performance: a 7B OmniAgent outperforms the 10x larger Qwen2.5-VL-72B on LVBench (50.5% vs. 47.3%).

Overview

  • Field: Computer Vision
  • Authors: Zhenghao Xing, Ruiyang Xu, Yuxuan Wang
  • Published: 2026-06-19
  • arXiv: 2506.14987

Summary

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly regardless of query difficulty, causing computational cost to grow with video duration. Although interactive frameworks have emerged, they often rely on global pre-scanning, and their context cost still scales with video length.

OmniAgent, the first native omni-modal agent, formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle. OmniAgent executes on-demand actions to selectively distill audio-visual cues into a persistent textual memory, effectively decoupling reasoning complexity from raw video duration.

Key Components

1. Agentic Supervised Fine-Tuning: Bootstraps native active perception via best-of-N trajectory synthesis with dual-stage quality control. 2. Agentic Reinforcement Learning: Employs TAURA (Turn-aware Adaptive Uncertainty Rescaled Advantage), which leverages turn-level entropy to steer credit assignment toward pivotal discovery turns.

Results

Crucially, OmniAgent exhibits positive test-time scaling—performance improves as the number of reasoning turns increases, validating the efficacy of active perception. Empirical results across ten benchmarks (e.g., VideoMME, LVBench) demonstrate state-of-the-art performance among open-source models. Notably, on LVBench, the 7B agent outperforms the 10x larger Qwen2.5-VL-72B (50.5% vs. 47.3%).

---

*Auto-collected on 2026-06-19*

Tags

#long-video-understanding#multimodal-agents#reinforcement-learning#pomdp#test-time-scaling#video-qa#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981506