跳到主要导航 跳到搜索 跳到主要内容

ACTIVE: OFFLINE REINFORCEMENT LEARNING VIA ADAPTIVE IMITATION AND IN-SAMPLE V -ENSEMBLE

  • Beihang University
  • Zhongguancun Laboratory

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Offline reinforcement learning (RL) aims to learn from static datasets and thus faces the challenge of value estimation errors for out-of-distribution actions. The in-sample learning scheme addresses this issue by performing implicit TD backups that does not query the values of unseen actions. However, pre-existing in-sample value learning and policy extraction methods suffer from over-regularization, limiting their performance on suboptimal or compositional datasets. In this paper, we analyze key factors in in-sample learning that might potentially hinder the use of a milder constraint. We propose Actor-Critic with Temperature adjustment and In-sample Value Ensemble (ACTIVE), a novel in-sample offline RL algorithm that leverages an ensemble of V -functions for critic training and adaptively adjusts the constraint level using dual gradient descent. We theoretically show that the V -ensemble suppresses the accumulation of initial value errors, thereby mitigating overestimation. Our experiments on the D4RL benchmarks demonstrate that ACTIVE alleviates overfitting of value functions and outperforms existing in-sample methods in terms of learning stability and policy optimality.

源语言英语
主期刊名13th International Conference on Learning Representations, ICLR 2025
出版商International Conference on Learning Representations, ICLR
71369-71387
页数19
ISBN(电子版)9798331320850
出版状态已出版 - 2025
活动13th International Conference on Learning Representations, ICLR 2025 - Singapore, 新加坡
期限: 24 4月 202528 4月 2025

出版系列

姓名13th International Conference on Learning Representations, ICLR 2025

会议

会议13th International Conference on Learning Representations, ICLR 2025
国家/地区新加坡
Singapore
时期24/04/2528/04/25

学术指纹

探究 'ACTIVE: OFFLINE REINFORCEMENT LEARNING VIA ADAPTIVE IMITATION AND IN-SAMPLE V -ENSEMBLE' 的科研主题。它们共同构成独一无二的学术指纹。

引用此