TY - JOUR
T1 - Online Human Behavior Learning via Dynamic Regressor Extension and Mixing With Fixed-Time Convergence
AU - Lin, Jie
AU - Wu, Huai Ning
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2025
Y1 - 2025
N2 - To improve the hybrid augmented intelligence of a human-in-the-loop (HiTL) control system, it is desirable to investigate the issue of human behavior learning (HBL), i.e., empower the machine to understand how a human expert performs a manipulation task. The human expert is commonly modeled as an optimal controller with unknown weighting matrices that depict the tradeoff between different control objectives. Therefore, the goal of this article is to determine the weighting matrices of the human objective function with fast convergence rate, which is usually pursued to achieve better efficiency and performance in practice. Accordingly, for a class of HiTL system, we propose a novel adaptive inverse optimal control (IOC) approach for online learning human behavior with a fixed-time guarantee. Our proposed method consists of two parts. In the first step, a dynamic regressor extension and mixing (DREM)-based estimation method is used for online learning of the human feedback gain with fixed-time convergence using the demonstrated system state measurement only. Then, with the estimated human feedback gain, a semidefinite programming (SDP) problem is solved to determine the weighting matrices of the human objective function. The simulation and the experiment on a steering control have validated the effectiveness and applicability of the developed approach.
AB - To improve the hybrid augmented intelligence of a human-in-the-loop (HiTL) control system, it is desirable to investigate the issue of human behavior learning (HBL), i.e., empower the machine to understand how a human expert performs a manipulation task. The human expert is commonly modeled as an optimal controller with unknown weighting matrices that depict the tradeoff between different control objectives. Therefore, the goal of this article is to determine the weighting matrices of the human objective function with fast convergence rate, which is usually pursued to achieve better efficiency and performance in practice. Accordingly, for a class of HiTL system, we propose a novel adaptive inverse optimal control (IOC) approach for online learning human behavior with a fixed-time guarantee. Our proposed method consists of two parts. In the first step, a dynamic regressor extension and mixing (DREM)-based estimation method is used for online learning of the human feedback gain with fixed-time convergence using the demonstrated system state measurement only. Then, with the estimated human feedback gain, a semidefinite programming (SDP) problem is solved to determine the weighting matrices of the human objective function. The simulation and the experiment on a steering control have validated the effectiveness and applicability of the developed approach.
KW - Dynamic regressor extension and mixing (DREM)
KW - fixed-time convergence
KW - human behavior learning (HBL)
KW - inverse optimal control (IOC)
UR - https://www.scopus.com/pages/publications/85209098567
U2 - 10.1109/TII.2024.3485814
DO - 10.1109/TII.2024.3485814
M3 - 文章
AN - SCOPUS:85209098567
SN - 1551-3203
VL - 21
SP - 1764
EP - 1772
JO - IEEE Transactions on Industrial Informatics
JF - IEEE Transactions on Industrial Informatics
IS - 2
ER -