Skip to main navigation Skip to search Skip to main content

Hierarchical Task Reasoning for Aerial Vision-and-Language Navigation

  • Beihang University

Research output: Contribution to journalArticlepeer-review

Abstract

Vision-and-Language Navigation (VLN) aims to enable embodied agents to navigate in complex 3D environments following natural language instructions. Compared to ground-based VLN, aerial VLN is challenged by more complex instructions and visual scenes, which lead to significant semantic ambiguity and limit the performance of traditional global alignment methods. To address this, we propose a Hierarchical Task Reasoning (HTR) framework that introduces explicit structure into the navigation decision process. Specifically, our HTR framework employs a Large Language Model (LLM) to break down a complex instruction into a structured hierarchy of subtasks, components, and primitive elements. Guided by this structured hierarchy, our approach first aligns with high-level subtasks for contextual reasoning, and then decomposes the execution into parallel streams for action and landmark refinement. The landmark stream utilizes slot attention to ground specific visual cues, clearly distinguishing between actions and visual references. This coarse-to-fine process resolves semantic ambiguity, leading to more precise navigation. Extensive experiments conducted on the public AerialVLN dataset demonstrate that our proposed method achieves competitive performance compared with current state-of-the-art approaches, with clear advantages in Success weighted by normalized Dynamic Time Warping (SDTW), achieving 13.1% on the AerialVLN-S dataset with a 2.9% absolute improvement.

Original languageEnglish
JournalIEEE Transactions on Vehicular Technology
DOIs
StateAccepted/In press - 2026

Keywords

  • Large Language Models (LLMs)
  • Slot Attention
  • Vision-and-Language Navigation (VLN)

Fingerprint

Dive into the research topics of 'Hierarchical Task Reasoning for Aerial Vision-and-Language Navigation'. Together they form a unique fingerprint.

Cite this