跳到主要导航 跳到搜索 跳到主要内容

Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models

  • Beihang University
  • National Key Laboratory of Complex System Control and Intelligent Agent Cooperation

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Diffusion models (DMs) are advanced generative models, yet recent research reveals their vulnerability to backdoor attacks, which establish hidden associations between input patterns and targeted model behavior, potentially causing malicious outputs during inference. These attacks pose significant risks, including model owner reputation damage and harmful content generation. However, existing defense methods often fail against backdoor attacks on diffusion models. To address this gap, we propose Diff-Cleanse, a two-stage defense framework. The first stage introduces a novel trigger inversion method for backdoor detection, and the second stage applies a structural pruning-based method for backdoor removal. Experiments on 373 models poisoned by three state-of-the-art attacks show that Diff-Cleanse achieves > 97% detection accuracy, completely remove backdoors and maintains the models' benign performance. Code is available at https://github.com/shymuel/diff-cleanse.

源语言英语
主期刊名2025 IEEE International Conference on Multimedia and Expo
主期刊副标题Journey to the Center of Machine Imagination, ICME 2025 - Conference Proceedings
出版商IEEE Computer Society
ISBN(电子版)9798331594954
DOI
出版状态已出版 - 2025
活动2025 IEEE International Conference on Multimedia and Expo, ICME 2025 - Nantes, 法国
期限: 30 6月 20254 7月 2025

出版系列

姓名Proceedings - IEEE International Conference on Multimedia and Expo
ISSN(印刷版)1945-7871
ISSN(电子版)1945-788X

会议

会议2025 IEEE International Conference on Multimedia and Expo, ICME 2025
国家/地区法国
Nantes
时期30/06/254/07/25

学术指纹

探究 'Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此