跳到主要导航 跳到搜索 跳到主要内容

VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models

  • Jiawei Liang
  • , Siyuan Liang*
  • , Aishan Liu
  • , Xiaochun Cao*
  • *此作品的通讯作者
  • Sun Yat-Sen University
  • National University of Singapore
  • Key Laboratory of Precision Opto-Mechatronics Technology (Ministry of Education)

科研成果: 期刊稿件文章同行评审

摘要

Autoregressive Visual Language Models (VLMs) demonstrate remarkable few-shot learning capabilities within a multimodal context. Recently, multimodal instruction tuning has emerged as a technique to further refine instruction-following abilities. However, we uncover the potential threat posed by backdoor attacks on autoregressive VLMs during instruction tuning. Adversaries can implant a backdoor by inserting poisoned samples with triggers embedded in instructions or images to datasets, enabling malicious manipulation of the victim model’s predictions with predefined triggers. However, the frozen visual encoder in autoregressive VLMs imposes constraints on learning conventional image triggers. Additionally, adversaries may lack access to the parameters and architectures of the victim model. To overcome these challenges, we introduce a multimodal instruction backdoor attack, namely VL-Trojan. Our approach facilitates image trigger learning through active reshaping of poisoned features and enhances black-box attack efficacy through an iterative character-level text trigger generation method. Our attack successfully induces target output during inference, significantly outperforming baselines (+15.68%) in ASR. Furthermore, our attack demonstrates robustness across various model scales, architectures and few-shot in-context reasoning scenarios. Our codes are available at https://github.com/JWLiang007/VL-Trojan.

源语言英语
页(从-至)3994-4013
页数20
期刊International Journal of Computer Vision
133
7
DOI
出版状态已出版 - 7月 2025

指纹

探究 'VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models' 的科研主题。它们共同构成独一无二的指纹。

引用此