跳到主要导航 跳到搜索 跳到主要内容

Cross-modality neighbor constraints based unbalanced multi-view text–image re-identification

  • Beihang University
  • Army Academy of Armored Forces

科研成果: 期刊稿件文章同行评审

摘要

Text-to-image Person Re-Identification (TIReID) is an identity retrieval task between visual and textual modalities. Previous research focuses on learning rich and diverse modality-shared semantic features and achieving excellent performance. However, they still have several notable limitations: (1)Noisy influence: Due to the difficulty of cross-modality annotation and the uncertainty of crowdsourced label quality, it is inevitable to introduce noisy labels by incorrect text–image pairs. (2)Sample imbalance: Datasets collected from real-world sources often face an unbalanced distribution of samples across different categories, which results in inconsistent parameters update progress with the training phrase. To address these issues, we propose a two-stage training pipeline for TIReID learning with noisy correspondence. Firstly, we employ a Noisy Correspondence Detector based on heterogeneous relation retrieval estimating confidence weights from each sample pair. Secondly, we design a multi-view triplet loss function, which leverages sample-level features to interact with global class centers, addressing sample imbalance and facilitating a smoother distribution in feature space. Finally, we utilize these clean samples to train the model through a progressive learning process. Extensive experiments on RSTPReid, CUHK-PEDES, and ICFG-PEDES demonstrate the effectiveness of our method against the state-of-the-art TIReID methods.

源语言英语
文章编号338
期刊Multimedia Systems
30
6
DOI
出版状态已出版 - 12月 2024

学术指纹

探究 'Cross-modality neighbor constraints based unbalanced multi-view text–image re-identification' 的科研主题。它们共同构成独一无二的学术指纹。

引用此