Skip to main navigation Skip to search Skip to main content

CSDNet: Contrastive Similarity Distillation Network for Multi-lingual Image-Text Retrieval

  • Shichen Lu
  • , Longteng Guo
  • , Xingjian He
  • , Xinxin Zhu
  • , Jing Liu
  • , Si Liu*
  • *Corresponding author for this work
  • Beihang University
  • CAS - Institute of Automation

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Cross-modal image-text retrieval is a crucial task in the field of vision and language, aimed at retrieving the relevant samples from one modality as per the given user expressed in another modality. While most methods developed for this task have focused on English, recent advances expanded the scope of this task to the Multi-lingual domain. However, these methods face challenges due to the limited availability of annotated data in non-English languages. In this work, we propose a novel method that leverages an English pre-training model as a teacher to improve Multi-lingual image-text retrieval performance. Our method trains a student model that produces better Multi-lingual image-text similarity scores by learning from the English image-text similarity scores of the trained teacher. We introduce the contrastive loss to align the two different representations of the image and text, and the Contrastive Similarity Distillation loss to align the Multi-lingual image-text distribution of the student with that of the English teacher. We evaluate our method on two popular datasets, i.e., MS-COCO and Flickr-30K, and achieve state-of-the-art performance. Our approach shows significant improvement over existing methods and has potential for practical applications.

Original languageEnglish
Title of host publicationImage and Graphics - 12th International Conference, ICIG 2023, Proceedings
EditorsHuchuan Lu, Risheng Liu, Wanli Ouyang, Hui Huang, Jiwen Lu, Jing Dong, Min Xu
PublisherSpringer Science and Business Media Deutschland GmbH
Pages385-395
Number of pages11
ISBN (Print)9783031463105
DOIs
StatePublished - 2023
Event12th International Conference on Image and Graphics, ICIG 2023 - Nanjing, China
Duration: 22 Sep 202324 Sep 2023

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume14357 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference12th International Conference on Image and Graphics, ICIG 2023
Country/TerritoryChina
CityNanjing
Period22/09/2324/09/23

Keywords

  • Image-Text Retrieval
  • Knowledge Distillation
  • Multi-Lingual

Fingerprint

Dive into the research topics of 'CSDNet: Contrastive Similarity Distillation Network for Multi-lingual Image-Text Retrieval'. Together they form a unique fingerprint.

Cite this