论文概要
研究领域: CV
作者: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
发布时间: 2026-08-27
arXiv: 2608.27456
中文摘要
多模态大语言模型(MLLMs)可以解读街景图像,但城市智能体的核心能力在于:当智能体开始移动后,这种局部感知证据是否仍然有用。本文研究了当前MLLM智能体能在多大程度上将局部城市感知转化为复杂真实规模城市中的可靠行动。我们提出了UrbanGround,首个基于香港全境3D地理空间数据构建的物理约束沙盒环境,支持第一人称视角的闭环交互和导航地图。智能体可直接进入3D城市并以第一人称探索。研究通过三个研究问题追踪空间问题的增长:首先测试智能体能否通过主动观察充分理解局部场景以回答空间问题;然后检验这种理解能否支持导航,随着目的地越来越远、越来越不明确;最后考察行为在路线可用性和行人运动变化时是否仍然有效。当代MLLM智能体在视觉识别和短程空间推理方面表现出有用的原子能力,但定向和行人感知移动仍不可靠。核心失败出现在长时间探索中:局部能力无法组合成持续的目标导向行为,错误会累积而无法有效纠正。
原文摘要
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent ...
自动采集于 2026-08-30
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。