Abstract
The end of the COVID-19 pandemic has reshaped epidemic screening and public health management, prompting a reassessment of advanced rapid testing techniques and community-wide implementation (i.e., mass testing) to track disease transmission and dynamically optimize epidemic control strategies. However, given resources and funding constraints, the joint optimization of control efforts with mass testing over time remains an open challenge. In this paper, we address a joint optimization problem of resource allocation for epidemic control and mass testing in light of evolving epidemics, which are characterized by temporarily shifting system dynamics and mixed observability of system states. To tackle this challenge, we propose a novel non-stationary Markov decision process (MDP) framework that dynamically incorporates newly acquired observational data from test results and refines control and testing strategies accordingly. We propose an efficient continual reinforcement learning (RL) approach to solve this challenging MDP within an offline-to-online paradigm. Our approach begins with offline pre-training of neural networks and subsequently fine-tunes them using newly acquired data for adaptive decision-making. The hallmark of our approach lies in the parallel learning and memory mechanism, which enables the RL agent to develop a generalized understanding of system dynamics across diverse environments, thereby improving solution optimality and robustness compared to state-of-the-art methods. This online learning-while-optimization approach broadens the applications of RL in solving MDPs with scarce historical data. Numerical experiments demonstrate the superiority of our proposed approach in parameter identifiability, as well as solution optimality and robustness under varying observability and dynamic patterns.
| Original language | English |
|---|---|
| Journal | IISE Transactions |
| DOIs | |
| State | Accepted/In press - 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- continual reinforcement learning
- Data-driven decision-making
- epidemic control
- MDP with mixed observability
- non-stationary MDP
Fingerprint
Dive into the research topics of 'Dynamic epidemic control with continual learning via mass testing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver