摘要
When future information of a mobile user such as trajectory is known, predictive resource allocation for video on-demand service can reduce energy consumption of base station or increase network throughput with ensured user experience. Traditional methods for predictive resource allocation first predict user information (say trajectory) and then optimize resource (say power) allocation. However, the prediction accuracy degrades as the prediction horizon increases. To deal with this issue, several recent works employed deep reinforcement learning for online decision-making by formulating the predictive resource allocation problem as Markov decision process (MDP). However, for this kind of MDP problems that is appropriately solved by reinforcement learning, existing works design the state in a trial-and-error manner. For constrained optimization problems, most existing reinforcement learning methods for wireless problems add penalty terms to the reward function with manually adjustable hyper-parameters to satisfy the constraints. This paper proposes an unsupervised deep learning method for online predictive resource allocation in an end-to-end manner, which can jointly predict information and optimize resource allocation. The proposed method is able to improve the performance of predictive resource allocation by online end-to-end unsupervised deep learning, and can systematically design the state of MDP and satisfy complex constraints such that the tedious trial-and-error methods for designing state and satisfying constraints are no longer necessary. We analyze the relationship between the unsupervised deep learning and deep reinforcement learning. Simulation results show that the proposed method needs almost the same energy consumption as deep reinforcement learning with a simplified state design process, which verifies the theoretical analysis.
| 投稿的翻译标题 | Predictive resource allocation: unsupervised learning of Markov decision processes |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1983-2000 |
| 页数 | 18 |
| 期刊 | Scientia Sinica Informationis |
| 卷 | 54 |
| 期 | 8 |
| DOI | |
| 出版状态 | 已出版 - 2024 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 7 经济适用的清洁能源
关键词
- Markov decision process
- complex constraint
- deep reinforcement learning
- predictive resource allocation
- state design
- unsupervised deep learning
学术指纹
探究 '预测资源分配: 马尔可夫决策过程的无监督学习' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver