跳到主要导航 跳到搜索 跳到主要内容

预测资源分配: 马尔可夫决策过程的无监督学习

  • Jiajun Wu
  • , Jianyu Zhao
  • , Chengjian Sun
  • , Chenyang Yang*
  • *此作品的通讯作者
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

When future information of a mobile user such as trajectory is known, predictive resource allocation for video on-demand service can reduce energy consumption of base station or increase network throughput with ensured user experience. Traditional methods for predictive resource allocation first predict user information (say trajectory) and then optimize resource (say power) allocation. However, the prediction accuracy degrades as the prediction horizon increases. To deal with this issue, several recent works employed deep reinforcement learning for online decision-making by formulating the predictive resource allocation problem as Markov decision process (MDP). However, for this kind of MDP problems that is appropriately solved by reinforcement learning, existing works design the state in a trial-and-error manner. For constrained optimization problems, most existing reinforcement learning methods for wireless problems add penalty terms to the reward function with manually adjustable hyper-parameters to satisfy the constraints. This paper proposes an unsupervised deep learning method for online predictive resource allocation in an end-to-end manner, which can jointly predict information and optimize resource allocation. The proposed method is able to improve the performance of predictive resource allocation by online end-to-end unsupervised deep learning, and can systematically design the state of MDP and satisfy complex constraints such that the tedious trial-and-error methods for designing state and satisfying constraints are no longer necessary. We analyze the relationship between the unsupervised deep learning and deep reinforcement learning. Simulation results show that the proposed method needs almost the same energy consumption as deep reinforcement learning with a simplified state design process, which verifies the theoretical analysis.

投稿的翻译标题Predictive resource allocation: unsupervised learning of Markov decision processes
源语言繁体中文
页(从-至)1983-2000
页数18
期刊Scientia Sinica Informationis
54
8
DOI
出版状态已出版 - 2024

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 7 - 经济适用的清洁能源
    可持续发展目标 7 经济适用的清洁能源

关键词

  • Markov decision process
  • complex constraint
  • deep reinforcement learning
  • predictive resource allocation
  • state design
  • unsupervised deep learning

学术指纹

探究 '预测资源分配: 马尔可夫决策过程的无监督学习' 的科研主题。它们共同构成独一无二的学术指纹。

引用此