DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
*Note: The original post is a long-form, essay-style Chinese write-up of the paper. Below is a structured English rendition preserving its key arguments and data.*
> "Humans are unique not because we fly, but because we imagine ourselves flying." — Antonio Damasio (opening epigraph of the post)
Key points
- Problem: Aerial Vision-Language Navigation (VLN) asks a drone to follow instructions like "find the red-roofed church" using only onboard camera/microphone perception. It is far harder than ground navigation because of top-down flattened viewpoints, dynamic disturbances (wind, battery, propeller noise), partial observability, and linguistic ambiguity.
- Prior VLA models (RT-2, OpenVLA, Dream-VLA) map perception directly to actions, but suffer from: (1) very short historical memory, (2) short planning horizons, and (3) no reliable stopping criterion — drones circle near targets or stop too early.
- Many navigation memory mechanisms accidentally leak future information into current decisions ("causal contamination"), inflating benchmark scores without real navigation ability.
- DreamFly enforces that at any decision time *t*, only observations strictly earlier than *t* are used, via temporal masking (analogous to autoregressive attention applied along the time axis).
- A learned dynamic weighting mechanism emphasizes task-relevant history and fades irrelevant observations.
- The result is a model with genuine navigation strategy that generalizes better to unseen environments.
- One-shot planners (A*, RRT) fail in partially observable aerial settings.
- DreamFly uses a diffusion model — which excels at multi-modal distributions — to generate a K-step future action sequence from the current state.
- Only the first action is executed; new visual feedback triggers re-planning (closed-loop receding-horizon control).
- The post likens this to a card game where you peek at the next 5 cards but may play only one.
- Traditional stopping is implicit — a byproduct of action generation — and interferes with planning.
- LiteStop estimates stop probability directly from action logits at the fully-masked (undecided) state, making termination an explicit, lightweight, independent decision.
- This reduces circling near targets and premature stops.
- Roughly +28% (seen) and +34% (unseen) relative success-rate improvements, with the lowest navigation error among compared methods — stronger gains on unseen maps indicate better generalization.
- Deng, Y., & Xu, F. (2026). *DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation*. arXiv:2608.12308.
- Driess et al. (2023). *PaLM-E*.
- Brohan et al. (2023). *RT-2*.
- Kim et al. (2025). *OpenVLA*.
- Song et al. (2025). *Dream-VLA*.
- Ho et al. (2020). *Denoising Diffusion Probabilistic Models*.
- Anderson et al. (2018). *Vision-and-Language Navigation*.
- Chen et al. (2024). *OpenFly: A Benchmark for Aerial Vision-Language Navigation*.
DreamFly's three components
1. Causally Aligned Historical Memory
2. Receding-Horizon Diffusion Planning (Plan-K, Execute-One)
3. LiteStop: Explicit Decoupled Termination
Results (OpenFly benchmark)
| Method | Test-Seen SR | Test-Unseen SR | Test-Seen SPL | Test-Unseen SPL | |---|---|---|---|---| | Prior best | ~25% | ~22% | ~20% | ~17% | | DreamFly | 32.04% | 29.46% | 28.22% | 23.54% |
Broader takeaways (from the post's conclusion)
1. Causal constraints are foundational for trustworthy AI: decisions must rely only on information genuinely available at decision time. 2. Rolling-horizon planning is a universal strategy for uncertainty in partially observable, dynamic environments (robot manipulation, autonomous driving). 3. Explicit termination — knowing when to stop as a decoupled decision — is a key capability for reliable, interpretable autonomous systems.
The post closes with the albatross metaphor: DreamFly translates evolved instincts — memory of past flight, short-term forecasting, and knowing when to dive — into mathematics for drones.