TY - JOUR
T1 - A meta-learning-guided TD3 control algorithm with adaptive experience replay for active vibration isolator
AU - Li, Weipeng
AU - Li, Haohui
AU - Ren, Xiaoguang
AU - Liu, Zeshu
AU - Cui, Yi
N1 - Publisher Copyright:
© 2026 Elsevier Ltd.
PY - 2026/9
Y1 - 2026/9
N2 - Structural vibration control is essential in aerospace, civil engineering, and electromechanical systems, yet the performance of conventional methods often degrades in nonlinear or time-varying environments owing to their strong dependence on accurate system models. Deep reinforcement learning (DRL) offers high adaptability and can effectively address these issues, but it typically suffers from low sample efficiency and unstable convergence. To overcome these limitations, this study proposes Meta-Twin Delayed Deep Deterministic Policy Gradient (Meta-TD3), a meta-learning-enhanced DRL framework. A Meta-Adaptive Experience Evaluator (MAEE) is introduced to assess transition utility through a nonlinear feature-to-priority mapping. Building on TD-error information, MAEE down-weights high-uncertainty transitions and maximizes entropy to preserve sampling diversity, enabling more reliable prioritization than PER-based methods. In the Hopper benchmark, Meta-TD3 achieves substantially faster convergence, requiring 52.3% fewer steps than TD3 and 17.6% fewer than PER-TD3, while also exhibiting improved training stability. For an active vibration isolation task on a single-degree-of-freedom platform, Meta-TD3 reduces peak vibration by 65% in simulations and 60% in experiments, and reduces RMS acceleration by 25.58% in simulations and 24.86% in experiments. Moreover, compared with the conventional Skyhook controller, Meta-TD3 delivers approximately 30% better suppression in fixed-frequency tests and enhances RMS reduction in sweep-frequency tests, demonstrating superior vibration attenuation under both operating conditions. These results demonstrate that Meta-TD3 enhances sample efficiency, improves stability, and provides a flexible and effective reinforcement learning-based framework for active vibration suppression in complex structural systems.
AB - Structural vibration control is essential in aerospace, civil engineering, and electromechanical systems, yet the performance of conventional methods often degrades in nonlinear or time-varying environments owing to their strong dependence on accurate system models. Deep reinforcement learning (DRL) offers high adaptability and can effectively address these issues, but it typically suffers from low sample efficiency and unstable convergence. To overcome these limitations, this study proposes Meta-Twin Delayed Deep Deterministic Policy Gradient (Meta-TD3), a meta-learning-enhanced DRL framework. A Meta-Adaptive Experience Evaluator (MAEE) is introduced to assess transition utility through a nonlinear feature-to-priority mapping. Building on TD-error information, MAEE down-weights high-uncertainty transitions and maximizes entropy to preserve sampling diversity, enabling more reliable prioritization than PER-based methods. In the Hopper benchmark, Meta-TD3 achieves substantially faster convergence, requiring 52.3% fewer steps than TD3 and 17.6% fewer than PER-TD3, while also exhibiting improved training stability. For an active vibration isolation task on a single-degree-of-freedom platform, Meta-TD3 reduces peak vibration by 65% in simulations and 60% in experiments, and reduces RMS acceleration by 25.58% in simulations and 24.86% in experiments. Moreover, compared with the conventional Skyhook controller, Meta-TD3 delivers approximately 30% better suppression in fixed-frequency tests and enhances RMS reduction in sweep-frequency tests, demonstrating superior vibration attenuation under both operating conditions. These results demonstrate that Meta-TD3 enhances sample efficiency, improves stability, and provides a flexible and effective reinforcement learning-based framework for active vibration suppression in complex structural systems.
KW - Deep reinforcement learning
KW - Meta-Twin delayed deep deterministic policy gradient
KW - Meta-adaptive experience evaluator
KW - Sampling distribution
KW - Structural vibration control
UR - https://www.scopus.com/pages/publications/105035640499
U2 - 10.1016/j.aei.2026.104672
DO - 10.1016/j.aei.2026.104672
M3 - 文章
AN - SCOPUS:105035640499
SN - 1474-0346
VL - 74
JO - Advanced Engineering Informatics
JF - Advanced Engineering Informatics
M1 - 104672
ER -