English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UrbanGround: Evaluating Spatial Agency of MLLM Agents in a Real-Scale 3D Replica of Hong Kong

Forum topic · 小凯 · 2026-08-30

Summary

UrbanGround is a physically constrained sandbox environment built from territory-wide 3D geospatial data of Hong Kong, designed to test whether multimodal large language model (MLLM) agents can convert local urban perception into reliable action at real-world scale. Agents explore the 3D city from a first-person view with closed-loop interaction and an interactive navigation map. The study traces how spatial demands grow across three research questions: understanding local scenes through active observation, navigating to increasingly distant and ambiguous destinations, and maintaining valid behavior amid route availability changes and pedestrian movement. Findings show that contemporary MLLM agents possess useful atomic capabilities in visual recognition and short-range spatial reasoning, but orientation and pedestrian-aware movement remain unreliable. The central failure emerges during long-horizon exploration, where local abilities fail to compose into sustained goal-directed behavior and errors accumulate without effective correction. Paper: arXiv 2608.27456.

Overview

  • Field: Computer Vision
  • arXiv: 2608.27456
  • Authors: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
  • Abstract

    Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. This paper investigates how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city.

    The authors propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person perspective.

    Three Research Questions

    1. Local scene understanding — Can an agent, through active observation, understand local scenes sufficiently well to answer spatial questions? 2. Navigation — Can this understanding support navigation as destinations become farther and more ambiguous? 3. Robust behavior — Does behavior remain valid when route availability and pedestrian movement change?

    Findings

  • Contemporary MLLM agents show useful atomic capabilities in visual recognition and short-range spatial reasoning.
  • Orientation and pedestrian-aware movement remain unreliable.
  • The core failure appears in long-horizon exploration: local capabilities fail to compose into sustained goal-directed behavior, and errors accumulate without effective correction.
---

*Auto-collected on 2026-08-30.*

Tags

#mlLM-agents#spatial-reasoning#urban-navigation#benchmarks#computer-vision#3d-environments#hong-kong#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634226