Skip to main navigation Skip to search Skip to main content

SDPT: Synchronous Dual Prompt Tuning for Fusion-Based Visual-Language Pre-trained Models

  • Yang Zhou
  • , Yongjian Wu
  • , Jiya Saiyin
  • , Bingzheng Wei
  • , Maode Lai
  • , Eric Chang
  • , Yan Xu*
  • *Corresponding author for this work
  • Beihang University
  • ByteDance Ltd.
  • Zhejiang University
  • Taiwan Artificial Intelligence Foundation

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Prompt tuning methods have achieved remarkable success in parameter-efficient fine-tuning on large pre-trained models. However, their application to dual-modal fusion-based visual-language pre-trained models (VLPMs), such as GLIP, has encountered issues. Existing prompt tuning methods have not effectively addressed the modal mapping and aligning problem for tokens in different modalities, leading to poor transfer generalization. To address this issue, we propose Synchronous Dual Prompt Tuning (SDPT). SDPT initializes a single set of learnable unified prototype tokens in the established modal aligning space to represent the aligned semantics of text and image modalities for downstream tasks. Furthermore, SDPT establishes inverse linear projections that require no training to embed the information of unified prototype tokens into the input space of different modalities. The inverse linear projections allow the unified prototype token to synchronously represent the two modalities and enable SDPT to share the unified semantics of text and image for downstream tasks across different modal prompts. Experimental results demonstrate that SDPT assists fusion-based VLPMs to achieve superior outcomes with only 0.04% of model parameters for training across various scenarios, outperforming other single- or dual-modal methods. The code will be released at https://github.com/wuyongjianCODE/SDPT.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages340-356
Number of pages17
ISBN (Print)9783031729669
DOIs
StatePublished - 2025
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: 29 Sep 20244 Oct 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15107 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period29/09/244/10/24

Keywords

  • Parameter-efficient fine-tuning
  • Prompt tuning
  • Visual-language pre-trained models

Fingerprint

Dive into the research topics of 'SDPT: Synchronous Dual Prompt Tuning for Fusion-Based Visual-Language Pre-trained Models'. Together they form a unique fingerprint.

Cite this