Paper Overview
Field: Computer Vision (CV) Authors: Wenxuan Guo, Xiuwei Xu, Yichen Liu Published: 2025-05-23 arXiv: 2505.17383
Abstract (translated from Chinese)
Vision-language navigation (VLN) requires agents to ground language instructions in their own movement within visual environments. Although state-of-the-art methods leverage the reasoning capabilities of vision-language models (VLMs) for end-to-end action prediction, they often lack an explicit and interpretable understanding of the relationships between the agent, the instruction, and the scene. In contrast, explicitly constructing scene maps for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large-scale vision-language pretraining.
To bridge this gap, the authors propose AwareVLN, a new framework that equips navigation models with a self-awareness reasoning mechanism, enabling them to understand agent state and task progress in a fully end-to-end and data-driven manner.
The method has two key innovations:
1. A structured reasoning module that fosters spatial and task-oriented self-awareness. 2. An automatic data engine with progress decomposition for efficient training.
Extensive experiments on multiple datasets in the Habitat simulator demonstrate that AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods.
Links
- Paper: https://arxiv.org/abs/2505.17383
*Auto-collected on 2026-05-23*