TY - GEN
T1 - Model-Free Zero Trust Defense Method Based on Recurrent Actor-Critic Framework
AU - Cheng, Yusi
AU - Sun, Jie
AU - Li, Bo
AU - Lu, Qihao
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Due to the complexity of network scenarios and inaccurate state observations, zero-trust defense (ZTD) architecture, which fundamentally operates as a defense mechanism based on user trustworthiness, is commonly modeled as partially observable markov decision process (POMDP). Existing control methods on zero-trust defense processes primarily evaluate trustworthiness in a model-based Bayesian manner, losing their generality. To overcome the difficulty of obtaining model knowledge in practical applications, this paper proposes a model-free reinforcement learning (RL) approach based on recurrent actorcritic (RAC) framework as the zero-trust policy engine. This method does not rely on observation matrix or state transition information. Instead, it uses a recurrent neural network (RNN) to assess trustworthiness from records of past observations, which is more aligned with real-world scenarios characterized by limited information and unknown attack types. To validate the effectiveness of the proposed method, we simulate a zerotrust attack-defense scenario by combining security datasets with real-world data. The proposed methods achieve excellent control performance in the simulated environment.
AB - Due to the complexity of network scenarios and inaccurate state observations, zero-trust defense (ZTD) architecture, which fundamentally operates as a defense mechanism based on user trustworthiness, is commonly modeled as partially observable markov decision process (POMDP). Existing control methods on zero-trust defense processes primarily evaluate trustworthiness in a model-based Bayesian manner, losing their generality. To overcome the difficulty of obtaining model knowledge in practical applications, this paper proposes a model-free reinforcement learning (RL) approach based on recurrent actorcritic (RAC) framework as the zero-trust policy engine. This method does not rely on observation matrix or state transition information. Instead, it uses a recurrent neural network (RNN) to assess trustworthiness from records of past observations, which is more aligned with real-world scenarios characterized by limited information and unknown attack types. To validate the effectiveness of the proposed method, we simulate a zerotrust attack-defense scenario by combining security datasets with real-world data. The proposed methods achieve excellent control performance in the simulated environment.
KW - model-free
KW - partially observable markov decision process
KW - recurrent actor-critic
KW - reinforcement learning
KW - zero-trust defense
UR - https://www.scopus.com/pages/publications/105030323380
U2 - 10.1109/CSCloud66326.2025.00057
DO - 10.1109/CSCloud66326.2025.00057
M3 - 会议稿件
AN - SCOPUS:105030323380
T3 - Proceedings - 2025 IEEE 12th International Conference on Cyber Security and Cloud Computing, CSCloud 2025
SP - 320
EP - 325
BT - Proceedings - 2025 IEEE 12th International Conference on Cyber Security and Cloud Computing, CSCloud 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 12th IEEE International Conference on Cyber Security and Cloud Computing, CSCloud 2025
Y2 - 7 November 2025 through 9 November 2025
ER -