English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

FineCog-Nav: Fine-Grained Cognitive Modules for Zero-Shot UAV Vision-Language Navigation

Forum topic · 小凯 · 2026-04-21

Summary

FineCog-Nav is a zero-shot framework for UAV vision-language navigation (VLN) inspired by human cognition. Instead of relying on a single large foundation model with generic prompts, it decomposes navigation into fine-grained cognitive modules—language processing, perception, attention, memory, imagination, reasoning, and decision-making—each driven by a moderate-sized foundation model with role-specific prompts and structured input-output protocols. This design enables effective module collaboration and improved interpretability. The authors also introduce AerialVLN-Fine, a benchmark of 300 curated trajectories from AerialVLN with sentence-level instruction-trajectory alignment and refined instructions containing explicit visual endpoints and landmark references. Experiments show FineCog-Nav consistently outperforms zero-shot baselines in instruction following, long-horizon planning, and generalization to unseen environments, demonstrating the value of fine-grained cognitive modularization for zero-shot aerial navigation. Paper: arXiv 2604.16298.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Dian Shao, Zhengzheng Xu, Peiyang Wang, Like Liu, Yule Wang, Jieqi Shi, Jing Huo
  • Published: 2026-04-17
  • arXiv: 2604.16298
  • Project page: https://smartdianlab.github.io/projects-FineCogNav
  • Summary

    UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous, multi-step instructions over long horizons. Existing zero-shot methods remain limited, as they often rely on large base models, generic prompts, and loosely coordinated modules.

    Key Contributions

  • FineCog-Nav framework: A top-down framework inspired by human cognition that organizes navigation into fine-grained modules for language processing, perception, attention, memory, imagination, reasoning, and decision-making.
  • Module design: Each module is driven by a moderate-sized foundation model with role-specific prompts and structured input-output protocols, enabling effective collaboration and improved interpretability.
  • AerialVLN-Fine benchmark: A fine-grained evaluation benchmark of 300 curated trajectories from AerialVLN, featuring sentence-level instruction-trajectory alignment and refined instructions with explicit visual endpoints and landmark references.

Results

Experiments show that FineCog-Nav consistently outperforms zero-shot baselines in instruction following, long-horizon planning, and generalization to unseen environments. These results demonstrate the effectiveness of fine-grained cognitive modularization for zero-shot aerial navigation.

---

*Auto-collected on 2026-04-21.*

Tags

#vision-language-navigation#uav#zero-shot#cognitive-architecture#embodied-ai#computer-vision#benchmark#aerial-navigation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618601