TY - GEN
T1 - Diff-Cleanse
T2 - 2025 IEEE International Conference on Multimedia and Expo, ICME 2025
AU - Jiang, Hao
AU - Xiao, Jin
AU - Hu, Xiaoguang
AU - Chen, Tianyou
AU - Zhao, Jiajia
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Diffusion models (DMs) are advanced generative models, yet recent research reveals their vulnerability to backdoor attacks, which establish hidden associations between input patterns and targeted model behavior, potentially causing malicious outputs during inference. These attacks pose significant risks, including model owner reputation damage and harmful content generation. However, existing defense methods often fail against backdoor attacks on diffusion models. To address this gap, we propose Diff-Cleanse, a two-stage defense framework. The first stage introduces a novel trigger inversion method for backdoor detection, and the second stage applies a structural pruning-based method for backdoor removal. Experiments on 373 models poisoned by three state-of-the-art attacks show that Diff-Cleanse achieves > 97% detection accuracy, completely remove backdoors and maintains the models' benign performance. Code is available at https://github.com/shymuel/diff-cleanse.
AB - Diffusion models (DMs) are advanced generative models, yet recent research reveals their vulnerability to backdoor attacks, which establish hidden associations between input patterns and targeted model behavior, potentially causing malicious outputs during inference. These attacks pose significant risks, including model owner reputation damage and harmful content generation. However, existing defense methods often fail against backdoor attacks on diffusion models. To address this gap, we propose Diff-Cleanse, a two-stage defense framework. The first stage introduces a novel trigger inversion method for backdoor detection, and the second stage applies a structural pruning-based method for backdoor removal. Experiments on 373 models poisoned by three state-of-the-art attacks show that Diff-Cleanse achieves > 97% detection accuracy, completely remove backdoors and maintains the models' benign performance. Code is available at https://github.com/shymuel/diff-cleanse.
KW - backdoor defense
KW - diffusion model security
KW - structural pruning
KW - trigger inversion
UR - https://www.scopus.com/pages/publications/105022607384
U2 - 10.1109/ICME59968.2025.11210014
DO - 10.1109/ICME59968.2025.11210014
M3 - 会议稿件
AN - SCOPUS:105022607384
T3 - Proceedings - IEEE International Conference on Multimedia and Expo
BT - 2025 IEEE International Conference on Multimedia and Expo
PB - IEEE Computer Society
Y2 - 30 June 2025 through 4 July 2025
ER -