TY - GEN
T1 - A Method for Robustness Testing of Intelligent Classification Models
T2 - 15th International Conference on Reliability, Maintenance and Safety, ICRMS 2024
AU - Liu, Yeyang
AU - Sun, Yangyang
AU - Wang, Yiwei
AU - Wu, Changjian
AU - Zhao, Changdi
AU - Jiang, Feng
AU - Ni, Liang
AU - Li, Xiaobin
AU - Yang, Dezhen
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Utilizing adversarial samples is essential for assessing the robustness of intelligent classification models. However, certain adversarial sample generation methods based on the grey-box approach face challenges, including low attack efficiency and limited interpretability. This paper presents the Score-based Gray-box Single-pixel Attacks (SGSA) method, a novel approach for generating adversarial samples. Initially, the feature map of the model's final layer is extracted, along with computing the discrepancy between the predicted score of each feature map and the original prediction score. By employing the Leaky ReLU activation function, the score of each pixel is computed. Subsequently, these scores are sorted in ascending order. Initiating the attack from the pixel with the lowest score, the most effective attack direction is determined by comparing the outcomes of both forward and reverse adversarial attacks on each pixel. The results demonstrate that this method not only achieves high attack efficiency but also offers interpretability in selecting attack locations, and the attack efficiency has doubled. Furthermore, this method offers valuable guidance for both testing and enhancing the model's robustness to some degree.
AB - Utilizing adversarial samples is essential for assessing the robustness of intelligent classification models. However, certain adversarial sample generation methods based on the grey-box approach face challenges, including low attack efficiency and limited interpretability. This paper presents the Score-based Gray-box Single-pixel Attacks (SGSA) method, a novel approach for generating adversarial samples. Initially, the feature map of the model's final layer is extracted, along with computing the discrepancy between the predicted score of each feature map and the original prediction score. By employing the Leaky ReLU activation function, the score of each pixel is computed. Subsequently, these scores are sorted in ascending order. Initiating the attack from the pixel with the lowest score, the most effective attack direction is determined by comparing the outcomes of both forward and reverse adversarial attacks on each pixel. The results demonstrate that this method not only achieves high attack efficiency but also offers interpretability in selecting attack locations, and the attack efficiency has doubled. Furthermore, this method offers valuable guidance for both testing and enhancing the model's robustness to some degree.
KW - adversarial attack
KW - adversarial sample
KW - gray-box
KW - robustness
UR - https://www.scopus.com/pages/publications/105030324815
U2 - 10.1109/ICRMS63553.2024.00170
DO - 10.1109/ICRMS63553.2024.00170
M3 - 会议稿件
AN - SCOPUS:105030324815
T3 - Proceedings - 2024 15th International Conference on Reliability, Maintenance and Safety, ICRMS 2024
SP - 1067
EP - 1073
BT - Proceedings - 2024 15th International Conference on Reliability, Maintenance and Safety, ICRMS 2024
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 31 July 2024 through 2 August 2024
ER -