Paper Overview
Field: cs.CV Authors: Like Liu, Zhengzheng Xu, Haitao He arXiv: 2507.15492Abstract
Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts.To address this gap, the authors present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective.
Key Contributions
- Dual-cognition evaluation: Includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, requiring spatial or temporal grounding beyond discrete answer prediction.
- Automated data pipeline: Constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples.
- Training split: UAV-DualCog-Train is built from disjoint scenes and shown via a lightweight optimization probe to provide useful structured supervision.
- Self-state reasoning
- Viewpoint transformation
- Precise spatial grounding
- Temporal interval localization
Findings
Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Persistent bottlenecks include:Conclusion
UAV-DualCog serves both as an evaluation benchmark and as a data resource for advancing MLLM-based UAV agents.--- *Originally posted on zhichai.net; auto-collected on 2026-07-21.*