摘要
The application of reinforcement learning (RL) in real-world scenarios is limited due to safety concerns and the distribution-shift challenge in the offline setting. To address these issues, we formulate safe offline RL as a constrained policy optimization problem that integrates cumulative cost constraints and behavioral policy regularization. We first derive the analytical solution of the offline constrained policy optimization problem through Lagrangian duality. Then, we prove that iterative updates of this solution guarantee monotonic performance improvement while bounding worst-case costs relative to the behavioral policy. To further prevent out-of-distribution actions that may violate safety constraints, we propose a mechanism that distills a “safe action” distribution from the offline data and restricts policy updates within this safe region. We term this approach safe anchoring. By projecting the analytical solution into a parameterized policy space using a VAE-distilled safe anchoring mechanism, we develop the Offline Constrained Policy Optimization with Safe Anchoring (OCPO-SA) algorithm. Extensive experiments on Safety-Gymnasium and Bullet-Safety-Gym demonstrate that OCPO-SA achieves safety in all tested environments, with the average cost reduced by 24% compared with the best-performing baseline among the compared algorithms.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 108865 |
| 期刊 | Neural Networks |
| 卷 | 201 |
| DOI | |
| 出版状态 | 已出版 - 9月 2026 |
学术指纹
探究 'Offline constrained policy optimization with safe anchoring' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver