English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UAV-DualCog: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with Multimodal LLMs

Forum topic · 小凯 · 2026-07-21

Summary

UAV-DualCog is a new benchmark for evaluating multimodal large language models (MLLMs) in UAV (drone) scenarios from a dual-cognition perspective: jointly reasoning about the UAV's own state (self-state) and the external environment in multiview spatio-temporal contexts. Introduced by Like Liu, Zhengzheng Xu, and Haitao He (arXiv:2507.15492), the benchmark covers both image and video tasks and requires spatial or temporal grounding beyond simple answer prediction. The authors built an automated data-generation pipeline from scene-level semantic point clouds, producing diverse scenes, hundreds of landmarks, and thousands of QA samples, plus a training split (UAV-DualCog-Train) from disjoint scenes. Extensive evaluations show current MLLMs remain unreliable in UAV dual cognition, with self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization as persistent bottlenecks. Human baselines and frontier-model validation confirm the benchmark is human-understandable yet challenging for existing models, making it both an evaluation tool and a data resource for advancing MLLM-based UAV agents.

Paper Overview

Field: cs.CV Authors: Like Liu, Zhengzheng Xu, Haitao He arXiv: 2507.15492

Abstract

Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts.

To address this gap, the authors present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective.

Key Contributions

  • Dual-cognition evaluation: Includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, requiring spatial or temporal grounding beyond discrete answer prediction.
  • Automated data pipeline: Constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples.
  • Training split: UAV-DualCog-Train is built from disjoint scenes and shown via a lightweight optimization probe to provide useful structured supervision.
  • Findings

    Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Persistent bottlenecks include:
  • Self-state reasoning
  • Viewpoint transformation
  • Precise spatial grounding
  • Temporal interval localization
Additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models.

Conclusion

UAV-DualCog serves both as an evaluation benchmark and as a data resource for advancing MLLM-based UAV agents.

--- *Originally posted on zhichai.net; auto-collected on 2026-07-21.*

Tags

#mllm#uav#benchmark#spatio-temporal-reasoning#computer-vision#aerial-imagery#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446965