跳到主要导航 跳到搜索 跳到主要内容

A Safety-Adjusted Policy Optimization Algorithm and Application for Obstacle Avoidance in the Quadcopter

  • Gang Xia
  • , Xinsong Yang*
  • , Qihan Qi
  • , Yaping Sun
  • , Xiwang Dong
  • *此作品的通讯作者
  • Sichuan University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Ensuring the safety of various real-world applications based on reinforcement learning (RL), such as quadcopter control, robotic manipulators, and autonomous robots, remains a critical challenge, despite RL's remarkable success in solving complex decision-making tasks. Existing on-policy Lagrangian optimization methods in safe RL typically use a single policy to balance the trade-off between safety and return without taking the potential benefits of adopting multiple policies into account. In this paper, a new on-policy method is proposed, named Safe-Adjusted Policy Optimization(SAPO), which is a dual-policy framework designed to address safety constraint violations in RL. By incorporating a cost-oriented policy to dynamically adjust a reward-oriented policy, the SAPO effectively resolves the trade-off between safety and return. Moreover, to enhance performance in carrying out high-dimensional tasks, the Kullback-Leibler (KL) divergence and a Gaussian kernel are employed in the distance functions to facilitate the training. In addition, a quadcopter-safe-navigation task is designed to overcome the drawback of previous research on quadcopter-safe-navigation with RL that only pays attention to reward function design without considering policy-level optimization. Finally, experimental results verify the feasibility of the designed task. Meanwhile, indicated by the test on real device, the proposed algorithm is easy to be implemented, offers performance guarantees, and outperforms existing safe RL baselines.

源语言英语
主期刊名IROS 2025 - 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, Conference Proceedings
编辑Christian Laugier, Alessandro Renzaglia, Nikolay Atanasov, Stan Birchfield, Grzegorz Cielniak, Leonardo De Mattos, Laura Fiorini, Philippe Giguere, Kenji Hashimoto, Javier Ibanez-Guzman, Tetsushi Kamegawa, Jinoh Lee, Giuseppe Loianno, Kevin Luck, Hisataka Maruyama, Philippe Martinet, Hadi Moradi, Urbano Nunes, Julien Pettre, Alberto Pretto, Tommaso Ranzani, Arne Ronnau, Silvia Rossi, Elliott Rouse, Fabio Ruggiero, Olivier Simonin, Danwei Wang, Ming Yang, Eiichi Yoshida, Huijing Zhao
出版商Institute of Electrical and Electronics Engineers Inc.
2360-2367
页数8
ISBN(电子版)9798331543938
DOI
出版状态已出版 - 2025
活动2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2025 - Hangzhou, 中国
期限: 19 10月 202525 10月 2025

出版系列

姓名IEEE International Conference on Intelligent Robots and Systems
ISSN(印刷版)2153-0858
ISSN(电子版)2153-0866

会议

会议2025 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2025
国家/地区中国
Hangzhou
时期19/10/2525/10/25

学术指纹

探究 'A Safety-Adjusted Policy Optimization Algorithm and Application for Obstacle Avoidance in the Quadcopter' 的科研主题。它们共同构成独一无二的学术指纹。

引用此