Skip to main navigation Skip to search Skip to main content

VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models

  • Jiawei Liang
  • , Siyuan Liang*
  • , Aishan Liu
  • , Xiaochun Cao*
  • *Corresponding author for this work
  • Sun Yat-Sen University
  • National University of Singapore
  • Key Laboratory of Precision Opto-Mechatronics Technology (Ministry of Education)

Research output: Contribution to journalArticlepeer-review

Abstract

Autoregressive Visual Language Models (VLMs) demonstrate remarkable few-shot learning capabilities within a multimodal context. Recently, multimodal instruction tuning has emerged as a technique to further refine instruction-following abilities. However, we uncover the potential threat posed by backdoor attacks on autoregressive VLMs during instruction tuning. Adversaries can implant a backdoor by inserting poisoned samples with triggers embedded in instructions or images to datasets, enabling malicious manipulation of the victim model’s predictions with predefined triggers. However, the frozen visual encoder in autoregressive VLMs imposes constraints on learning conventional image triggers. Additionally, adversaries may lack access to the parameters and architectures of the victim model. To overcome these challenges, we introduce a multimodal instruction backdoor attack, namely VL-Trojan. Our approach facilitates image trigger learning through active reshaping of poisoned features and enhances black-box attack efficacy through an iterative character-level text trigger generation method. Our attack successfully induces target output during inference, significantly outperforming baselines (+15.68%) in ASR. Furthermore, our attack demonstrates robustness across various model scales, architectures and few-shot in-context reasoning scenarios. Our codes are available at https://github.com/JWLiang007/VL-Trojan.

Original languageEnglish
Pages (from-to)3994-4013
Number of pages20
JournalInternational Journal of Computer Vision
Volume133
Issue number7
DOIs
StatePublished - Jul 2025

Keywords

  • Backdoor attack
  • Instruction tuning
  • Visual language model

Fingerprint

Dive into the research topics of 'VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models'. Together they form a unique fingerprint.

Cite this