English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

Forum topic · 小凯 · 2026-05-23

Summary

AwareVLN (arXiv:2505.17383) is a new vision-language navigation (VLN) framework that enables agents to reason with self-awareness when following natural-language instructions in visual environments. While state-of-the-art VLN methods use vision-language models for end-to-end action prediction, they often lack an explicit, interpretable understanding of relationships between the agent, the instruction, and the scene. Conversely, approaches that explicitly build scene maps for heuristic planning are intuitive but require extra 3D sensors and hinder large-scale vision-language pretraining. AwareVLN bridges this gap by equipping navigation models with a self-awareness reasoning mechanism, allowing fully end-to-end, data-driven understanding of agent state and task progress. Key innovations include: (1) a structured reasoning module that fosters spatial and task-oriented self-awareness, and (2) an automatic data engine with progress decomposition for efficient training. Experiments in the Habitat simulator across multiple datasets show AwareVLN significantly outperforms prior state-of-the-art VLN methods.

Paper Overview

Field: Computer Vision (CV) Authors: Wenxuan Guo, Xiuwei Xu, Yichen Liu Published: 2025-05-23 arXiv: 2505.17383

Abstract (translated from Chinese)

Vision-language navigation (VLN) requires agents to ground language instructions in their own movement within visual environments. Although state-of-the-art methods leverage the reasoning capabilities of vision-language models (VLMs) for end-to-end action prediction, they often lack an explicit and interpretable understanding of the relationships between the agent, the instruction, and the scene. In contrast, explicitly constructing scene maps for heuristic planning is intuitively appealing but relies on additional 3D sensors and hinders large-scale vision-language pretraining.

To bridge this gap, the authors propose AwareVLN, a new framework that equips navigation models with a self-awareness reasoning mechanism, enabling them to understand agent state and task progress in a fully end-to-end and data-driven manner.

The method has two key innovations:

1. A structured reasoning module that fosters spatial and task-oriented self-awareness. 2. An automatic data engine with progress decomposition for efficient training.

Extensive experiments on multiple datasets in the Habitat simulator demonstrate that AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods.

Links

  • Paper: https://arxiv.org/abs/2505.17383
---

*Auto-collected on 2026-05-23*

Tags

#vision-language-navigation#embodied-ai#computer-vision#vision-language-models#self-awareness-reasoning#habitat-simulator#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620659