TY - GEN
T1 - GIP
T2 - 32nd IEEE International Conference on Image Processing, ICIP 2025
AU - Lin, Xiang
AU - Li, Weixin
AU - Guo, Shu
AU - Wang, Lihong
AU - Huang, Di
N1 - Publisher Copyright:
©2025 IEEE.
PY - 2025
Y1 - 2025
N2 - —Existing Parameter Efficient Fine-Tuning (PEFT) methods in vision-language (VL) domains, primarily adapted from single-modality approaches, face limitations in modeling cross-modal interactions. These methods often unify visual and textual features without explicit modality-specific processing, or rely on unidirectional interaction, leading to suboptimal task adaptation for pre-trained Vision-Language Models (VLMs). To address this issue, we propose a Gated Interaction Prompt (GIP) module as a plug-and-play adaptation to existing PEFT methods, which effectively enhances the two-way interaction between visual and textual features. Our GIP module integrates learnable prompts alongside visual and textual features into the attention layers of VLMs, serving as a bridge for cross-modal interaction. Furthermore, GIP introduces task-specific gating mechanisms to regulate and adapt the influence of prompts across different tasks, thereby further enhancing model performance. Extensive experiments on four VL tasks demonstrate that our approach can seamlessly integrate with existing methods and achieves significant performance improvements with minimal impact on parameter counts and computational costs. With only a 0.02% increase in trainable parameters, our method achieves performance gains of 0.6%, 0.8%, and 1.2% across four tasks—when applied to VL-PET, VL-Adapter, and LoRA, respectively.
AB - —Existing Parameter Efficient Fine-Tuning (PEFT) methods in vision-language (VL) domains, primarily adapted from single-modality approaches, face limitations in modeling cross-modal interactions. These methods often unify visual and textual features without explicit modality-specific processing, or rely on unidirectional interaction, leading to suboptimal task adaptation for pre-trained Vision-Language Models (VLMs). To address this issue, we propose a Gated Interaction Prompt (GIP) module as a plug-and-play adaptation to existing PEFT methods, which effectively enhances the two-way interaction between visual and textual features. Our GIP module integrates learnable prompts alongside visual and textual features into the attention layers of VLMs, serving as a bridge for cross-modal interaction. Furthermore, GIP introduces task-specific gating mechanisms to regulate and adapt the influence of prompts across different tasks, thereby further enhancing model performance. Extensive experiments on four VL tasks demonstrate that our approach can seamlessly integrate with existing methods and achieves significant performance improvements with minimal impact on parameter counts and computational costs. With only a 0.02% increase in trainable parameters, our method achieves performance gains of 0.6%, 0.8%, and 1.2% across four tasks—when applied to VL-PET, VL-Adapter, and LoRA, respectively.
KW - Gating Mechanism
KW - Parameter-Efficient Fine-Tuning
KW - Pre-trained Vision-Language Models
UR - https://www.scopus.com/pages/publications/105028566230
U2 - 10.1109/ICIP55913.2025.11084554
DO - 10.1109/ICIP55913.2025.11084554
M3 - 会议稿件
AN - SCOPUS:105028566230
T3 - Proceedings - International Conference on Image Processing, ICIP
SP - 617
EP - 622
BT - 2025 IEEE International Conference on Image Processing, ICIP 2025 - Proceedings
PB - IEEE Computer Society
Y2 - 14 September 2025 through 17 September 2025
ER -