论文概要
Research Area: Computer Vision (CV) Authors: Yan Deng, Fei Xu Published: 2026-08-13 arXiv: 2508.03419
Introduction
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when a navigation goal has been reached under partial observability. While recent vision-language-action (VLA) models offer a promising perception-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination.
The DreamFly Framework
DreamFly is a diffusion-based aerial VLN framework built on Dream-VLA, with three key components:
1. Causal Aligned Historical Memory
DreamFly introduces causal alignment of historical memory, enhancing the current visual representation using only observations from before the current decision step. This enables temporal reasoning without leaking future information.2. Receding-Horizon Diffusion Planning
Navigation is formalized as receding-horizon diffusion planning: the policy predicts K-step action chunks but executes only the first action before replanning. This "plan K steps, execute one" strategy uses future actions as auxiliary planning objectives while preserving closed-loop visual feedback.3. LiteStop: Explicit Termination
LiteStop estimates the stop probability directly from the action logits starting from an initial fully-masked state, decoupling explicit termination from action generation.Experimental Results
Experiments on the OpenFly benchmark show consistent improvements in both seen and unseen environments:
| Split | SR | SPL | |-------|-----|-----| | Test Seen | 32.04% | 28.22% | | Test Unseen | 29.46% | 23.54% |
DreamFly outperforms all compared methods on both metrics while achieving the lowest navigation error.
Conclusion
These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.
Original Abstract
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time... (see the full paper on arXiv)
--- *Auto-collected on 2026-08-14*