跳到主要导航 跳到搜索 跳到主要内容

Offline constrained policy optimization with safe anchoring

  • Diyuan Hou
  • , Longyang Huang
  • , Pu Feng
  • , Wenjun Wu*
  • *此作品的通讯作者
  • Beihang University
  • Shanghai Jiao Tong University

科研成果: 期刊稿件文章同行评审

摘要

The application of reinforcement learning (RL) in real-world scenarios is limited due to safety concerns and the distribution-shift challenge in the offline setting. To address these issues, we formulate safe offline RL as a constrained policy optimization problem that integrates cumulative cost constraints and behavioral policy regularization. We first derive the analytical solution of the offline constrained policy optimization problem through Lagrangian duality. Then, we prove that iterative updates of this solution guarantee monotonic performance improvement while bounding worst-case costs relative to the behavioral policy. To further prevent out-of-distribution actions that may violate safety constraints, we propose a mechanism that distills a “safe action” distribution from the offline data and restricts policy updates within this safe region. We term this approach safe anchoring. By projecting the analytical solution into a parameterized policy space using a VAE-distilled safe anchoring mechanism, we develop the Offline Constrained Policy Optimization with Safe Anchoring (OCPO-SA) algorithm. Extensive experiments on Safety-Gymnasium and Bullet-Safety-Gym demonstrate that OCPO-SA achieves safety in all tested environments, with the average cost reduced by 24% compared with the best-performing baseline among the compared algorithms.

源语言英语
文章编号108865
期刊Neural Networks
201
DOI
出版状态已出版 - 9月 2026

学术指纹

探究 'Offline constrained policy optimization with safe anchoring' 的科研主题。它们共同构成独一无二的学术指纹。

引用此