English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

Forum topic · 小凯 · 2026-05-25

Summary

AwareVLN is a new framework for Vision-Language Navigation (VLN) that equips navigation models with self-awareness reasoning, enabling agents to understand their own state and task progress in a fully end-to-end, data-driven manner. While state-of-the-art VLN methods leverage vision-language models (VLMs) for end-to-end action prediction, they often lack an explicit, interpretable understanding of the relationships between the agent, the instruction, and the scene. Conversely, approaches that explicitly build scene maps for heuristic planning are intuitive but require additional 3D sensors and hinder large-scale vision-language pretraining. AwareVLN bridges this gap with two key innovations: (1) a structured reasoning module that promotes spatial and task-oriented self-awareness, and (2) an automatic data engine with progress decomposition for effective training. Extensive experiments across multiple datasets in the Habitat simulator show that AwareVLN significantly outperforms prior state-of-the-art VLN methods. Paper: arXiv 2505.14487.

Paper Overview

Research Area: Computer Vision (CV) Authors: Wenxuan Guo, Xiuwei Xu, Yichen Liu Published: 2026-05-25 arXiv: 2505.14487

Abstract

Vision-Language Navigation (VLN) requires agents to ground language instructions in their own egocentric observations within a visual environment. Although state-of-the-art methods leverage the reasoning capabilities of vision-language models (VLMs) for end-to-end action prediction, they often lack an explicit and interpretable understanding of the relationships between the agent, the instruction, and the scene. Conversely, explicitly constructing scene maps for heuristic planning is intuitively appealing, but relies on additional 3D sensors and hinders large-scale vision-language pretraining.

To bridge this gap, the authors propose AwareVLN, a novel framework that equips navigation models with a self-awareness reasoning mechanism, enabling them to understand agent state and task progress in a fully end-to-end and data-driven manner.

The method features two key innovations:

1. A structured reasoning module that facilitates spatial and task-oriented self-awareness. 2. An automatic data engine with progress decomposition to enable effective training.

Extensive experiments across multiple datasets in the Habitat simulator demonstrate that AwareVLN significantly outperforms previous state-of-the-art vision-language navigation methods.

--- *Auto-collected on 2026-05-25*

Tags

#vision-language-navigation#awarevln#vlm#robotics#embodied-ai#arxiv#cv#self-awareness

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620760