UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
- Research area: Computer Vision (CV)
- Authors: Tianjie Ju, Zheng Wu, Yueqing Sun
- Published: 2026-08-28
- arXiv: 2508.11373
- Current MLLM agents show useful atomic capabilities in visual recognition and short-range spatial reasoning.
- Direction awareness and pedestrian-aware motion remain unreliable.
- The core failure appears in long-horizon exploration: local capabilities do not compose into sustained goal-directed behavior, as errors accumulate without effective correction.
Summary
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. This paper investigates how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city.
The authors propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person perspective.
Analysis Structure
The analysis follows the growth of the spatial problem through three research questions:
1. Local grounding: Can the agent sufficiently ground local scenes after active observation to answer spatial questions? 2. Navigation: Does such grounding support navigation as destinations become farther and more ambiguous? 3. Robustness: Does the resulting behavior remain stable amid changes in route availability and pedestrian movement?
Key Findings
Original Abstract (excerpt)
> Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view...
*Auto-collected on 2026-08-29.*