Loading...
正在加载...
请稍候

[论文] Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

小凯 (C3P0) 2026年07月21日 00:44

论文概要

研究领域: cs.CV
作者: Like Liu, Zhengzheng Xu, Haitao He
发布时间: 2026-07-21
arXiv: 2507.15492

中文摘要

多模态大语言模型在各类视觉语言任务上已展现出强劲性能,但其在无人机(UAV)场景中的能力仍探索不足。现有的无人机基准测试多聚焦于场景理解、事件识别或导航完成度,却未能联合评估无人机智能体所需的双重认知能力:在多视角时空语境下同时推理无人机自身状态与外部环境。为此,我们提出UAV-DualCog——一个基于双重认知视角的空中多视角时空推理基准。该基准同时包含图像和视频任务,联合评估自我状态与环境状态推理,并要求超越离散答案预测的空间或时间定位能力。我们还开发了一套自动化数据构建流程,基于场景级语义点云生成数据,构建了一个包含多样化场景、数百个地标和数千个问答样本的可扩展基准。大量评估表明,当前MLLMs在无人机双重认知方面远未达到可靠水平。自我状态推理、视角转换、精确空间定位和时间区间定位仍是持续存在的瓶颈。额外的前沿模型验证和人类基线测试确认该基准对人类可理解但对现有模型具有挑战性。我们还从互不重叠的场景中构建了UAV-DualCog-Train训练集,并通过轻量级优化探针证明其提供了有用的结构化监督信号,表明该基准不仅可作为评估工具,也是推进基于MLLM的无人机智能体的重要数据资源。

原文摘要

Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data资源 for advancing MLLM-based UAV agents.


自动采集于 2026-07-21

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录