English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Forum topic · 小凯 · 2026-08-14

Summary

DreamFly is a diffusion-based aerial vision-language navigation (VLN) framework built on Dream-VLA, addressing three key limitations of VLA models in aerial navigation: limited historical context, short planning horizons, and unreliable implicit termination. The framework introduces a causally aligned historical memory that augments current visual representations using only observations from before the current decision step, enabling temporal reasoning without future information leakage. Navigation is formulated as receding-horizon diffusion planning, where the policy predicts K-step action chunks but executes only the first action before replanning—an approach that uses future actions as auxiliary planning objectives while retaining closed-loop visual feedback. Additionally, LiteStop estimates stopping probability directly from action logits starting from a fully masked state, decoupling explicit termination from action generation. On the OpenFly benchmark, DreamFly achieves 32.04%/29.46% success rate (SR) and 28.22%/23.54% SPL on seen/unseen splits, outperforming all baselines on both metrics with the lowest navigation error. The work is available on arXiv as 2508.03419.

论文概要

Research Area: Computer Vision (CV) Authors: Yan Deng, Fei Xu Published: 2026-08-13 arXiv: 2508.03419

Introduction

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when a navigation goal has been reached under partial observability. While recent vision-language-action (VLA) models offer a promising perception-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination.

The DreamFly Framework

DreamFly is a diffusion-based aerial VLN framework built on Dream-VLA, with three key components:

1. Causal Aligned Historical Memory

DreamFly introduces causal alignment of historical memory, enhancing the current visual representation using only observations from before the current decision step. This enables temporal reasoning without leaking future information.

2. Receding-Horizon Diffusion Planning

Navigation is formalized as receding-horizon diffusion planning: the policy predicts K-step action chunks but executes only the first action before replanning. This "plan K steps, execute one" strategy uses future actions as auxiliary planning objectives while preserving closed-loop visual feedback.

3. LiteStop: Explicit Termination

LiteStop estimates the stop probability directly from the action logits starting from an initial fully-masked state, decoupling explicit termination from action generation.

Experimental Results

Experiments on the OpenFly benchmark show consistent improvements in both seen and unseen environments:

| Split | SR | SPL | |-------|-----|-----| | Test Seen | 32.04% | 28.22% | | Test Unseen | 29.46% | 23.54% |

DreamFly outperforms all compared methods on both metrics while achieving the lowest navigation error.

Conclusion

These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.

Original Abstract

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time... (see the full paper on arXiv)

--- *Auto-collected on 2026-08-14*

Tags

#aerial-navigation#vision-language-navigation#vlm#diffusion-planning#embodied-ai#robotics#arxiv#paper-summary

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633453