Abstract
Multiple autonomous underwater vehicles (AUVs) target tracking problem is a significant challenge for AUV swarm control, which is crucial to the growth of the marine industry. To emphasize the great adaptability while tackling the limitations of reinforcement learning (RL) methods in Multi-AUV target tracking tasks, we propose an efficient two-stage learning from demonstrations (LfD) training framework, FISHER, based on few-shot expert demonstration, featuring imitation learning (IL) and offline reinforcement learning (ORL). In the first stage, we develop a sample-efficient algorithm, multi-agent discriminator actor-critic (MADAC), to facilitate the imitation of expert policy and the generation of offline datasets. In the second stage, based on the decision transformer (DT), the reward function-independent algorithm, multi-agent independent generalized decision transformer (MAIGDT) is utilized for further policy improvement. Simultaneously, we propose a simulation to simulation (sim2sim) method to facilitate the generation of expert trajectories, which is compatible with traditional methods like artificial potential field (APF). Through comparative experiments, we verify the improvement of the proposed MADAC and MAIGDT algorithms. Finally, full target tracking simulation processes show that FISHER can achkmieve performance comparable to expert demonstrations, thereby further demonstrating the strong practicality of FISHER framework. To accelerate relevant research in this direction, the code for simulation will be released as open-source.
| Original language | English |
|---|---|
| Title of host publication | Neural Information Processing - 31st International Conference, ICONIP 2024, Proceedings |
| Editors | Mufti Mahmud, Maryam Doborjeh, Zohreh Doborjeh, Kevin Wong, Andrew Chi Sing Leung, M. Tanveer |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 61-75 |
| Number of pages | 15 |
| ISBN (Print) | 9789819670352 |
| DOIs | |
| State | Published - 2026 |
| Event | 31st International Conference on Neural Information Processing, ICONIP 2024 - Auckland, New Zealand Duration: 2 Dec 2024 → 6 Dec 2024 |
Publication series
| Name | Communications in Computer and Information Science |
|---|---|
| Volume | 2297 CCIS |
| ISSN (Print) | 1865-0929 |
| ISSN (Electronic) | 1865-0937 |
Conference
| Conference | 31st International Conference on Neural Information Processing, ICONIP 2024 |
|---|---|
| Country/Territory | New Zealand |
| City | Auckland |
| Period | 2/12/24 → 6/12/24 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 14 Life Below Water
Keywords
- Autonomous underwater vehicles
- Learning from demonstrations
- Reinforcement learning
- Simulation to simulation
- Target tracking
Fingerprint
Dive into the research topics of 'FISHER: An Efficient Sim2sim Training Framework Dedicated in Multi-AUV Target Tracking via Learning from Demonstrations'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver