Skip to main navigation Skip to search Skip to main content

GIP: Gated Interaction Prompt for Parameter Efficient Vision-Language Fine-Tuning

  • Xiang Lin
  • , Weixin Li*
  • , Shu Guo
  • , Lihong Wang
  • , Di Huang
  • *Corresponding author for this work
  • Beihang University
  • Shanghai Artificial Intelligence Laboratory
  • National Computer Network Emergency Response Technical Team

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

—Existing Parameter Efficient Fine-Tuning (PEFT) methods in vision-language (VL) domains, primarily adapted from single-modality approaches, face limitations in modeling cross-modal interactions. These methods often unify visual and textual features without explicit modality-specific processing, or rely on unidirectional interaction, leading to suboptimal task adaptation for pre-trained Vision-Language Models (VLMs). To address this issue, we propose a Gated Interaction Prompt (GIP) module as a plug-and-play adaptation to existing PEFT methods, which effectively enhances the two-way interaction between visual and textual features. Our GIP module integrates learnable prompts alongside visual and textual features into the attention layers of VLMs, serving as a bridge for cross-modal interaction. Furthermore, GIP introduces task-specific gating mechanisms to regulate and adapt the influence of prompts across different tasks, thereby further enhancing model performance. Extensive experiments on four VL tasks demonstrate that our approach can seamlessly integrate with existing methods and achieves significant performance improvements with minimal impact on parameter counts and computational costs. With only a 0.02% increase in trainable parameters, our method achieves performance gains of 0.6%, 0.8%, and 1.2% across four tasks—when applied to VL-PET, VL-Adapter, and LoRA, respectively.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PublisherIEEE Computer Society
Pages617-622
Number of pages6
ISBN (Electronic)9798331523794
DOIs
StatePublished - 2025
Event32nd IEEE International Conference on Image Processing, ICIP 2025 - Anchorage, United States
Duration: 14 Sep 202517 Sep 2025

Publication series

NameProceedings - International Conference on Image Processing, ICIP
ISSN (Print)1522-4880

Conference

Conference32nd IEEE International Conference on Image Processing, ICIP 2025
Country/TerritoryUnited States
CityAnchorage
Period14/09/2517/09/25

Keywords

  • Gating Mechanism
  • Parameter-Efficient Fine-Tuning
  • Pre-trained Vision-Language Models

Fingerprint

Dive into the research topics of 'GIP: Gated Interaction Prompt for Parameter Efficient Vision-Language Fine-Tuning'. Together they form a unique fingerprint.

Cite this