跳到主要导航 跳到搜索 跳到主要内容

Model-free Reinforcement Learning with Stochastic Reward Stabilization for Recommender Systems

  • Tianchi Cai*
  • , Shenliao Bao
  • , Jiyan Jiang
  • , Shiji Zhou
  • , Wenpeng Zhang
  • , Lihong Gu
  • , Jinjie Gu
  • , Guannan Zhang
  • *此作品的通讯作者
  • Ant Group
  • Tsinghua University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Model-free RL-based recommender systems have recently received increasing research attention due to their capability to handle partial feedback and long-term rewards. However, most existing research has ignored a critical feature in recommender systems: one user's feedback on the same item at different times is random. The stochastic rewards property essentially differs from that in classic RL scenarios with deterministic rewards, which makes RL-based recommender systems much more challenging. In this paper, we first demonstrate in a simulator environment where using direct stochastic feedback results in a significant drop in performance. Then to handle the stochastic feedback more efficiently, we design two stochastic reward stabilization frameworks that replace the direct stochastic feedback with that learned by a supervised model. Both frameworks are model-agnostic, i.e., they can effectively utilize various supervised models. We demonstrate the superiority of the proposed frameworks over different RL-based recommendation baselines with extensive experiments on a recommendation simulator as well as an industrial-level recommender system.

源语言英语
主期刊名SIGIR 2023 - Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval
出版商Association for Computing Machinery, Inc
2179-2183
页数5
ISBN(电子版)9781450394086
DOI
出版状态已出版 - 18 7月 2023
已对外发布
活动46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023 - Taipei, 中国台湾
期限: 23 7月 202327 7月 2023

出版系列

姓名SIGIR 2023 - Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

会议

会议46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023
国家/地区中国台湾
Taipei
时期23/07/2327/07/23

指纹

探究 'Model-free Reinforcement Learning with Stochastic Reward Stabilization for Recommender Systems' 的科研主题。它们共同构成独一无二的指纹。

引用此