English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Replica

Forum topic · 小凯 · 2026-08-29

Summary

UrbanGround is a benchmark sandbox built from territory-wide 3D geospatial data of Hong Kong, designed to test whether multimodal large language model (MLLM) agents can convert local urban perception into reliable spatial action at real-world scale. The environment supports closed-loop, first-person exploration with an interactive navigation map. The authors structure their analysis around three research questions: (1) whether agents can sufficiently ground local scenes after active observation to answer spatial questions; (2) whether such grounding supports navigation as destinations become farther and more ambiguous; and (3) whether resulting behaviors remain stable under changing route availability and pedestrian movement. Findings show that current MLLM agents possess useful atomic capabilities in visual recognition and short-range spatial reasoning, but direction awareness and pedestrian-aware motion remain unreliable. The core failure emerges in long-horizon exploration: local abilities do not compose into sustained goal-directed behavior, as errors accumulate without effective correction. Paper: arXiv:2508.11373.

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

  • Research area: Computer Vision (CV)
  • Authors: Tianjie Ju, Zheng Wu, Yueqing Sun
  • Published: 2026-08-28
  • arXiv: 2508.11373
  • Summary

    Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. This paper investigates how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city.

    The authors propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person perspective.

    Analysis Structure

    The analysis follows the growth of the spatial problem through three research questions:

    1. Local grounding: Can the agent sufficiently ground local scenes after active observation to answer spatial questions? 2. Navigation: Does such grounding support navigation as destinations become farther and more ambiguous? 3. Robustness: Does the resulting behavior remain stable amid changes in route availability and pedestrian movement?

    Key Findings

  • Current MLLM agents show useful atomic capabilities in visual recognition and short-range spatial reasoning.
  • Direction awareness and pedestrian-aware motion remain unreliable.
  • The core failure appears in long-horizon exploration: local capabilities do not compose into sustained goal-directed behavior, as errors accumulate without effective correction.

Original Abstract (excerpt)

> Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view...

*Auto-collected on 2026-08-29.*

Tags

#mllm#embodied-agents#urban-navigation#computer-vision#benchmark#3d-environment#spatial-reasoning#hong-kong

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634187