TY - JOUR
T1 - Cross-modality neighbor constraints based unbalanced multi-view text–image re-identification
AU - Li, Yongxi
AU - Tang, Wenzhong
AU - Zhang, Ke
AU - Zhu, Xi
AU - Wang, Haoming
AU - Wang, Shuai
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer-Verlag GmbH Germany, part of Springer Nature 2024.
PY - 2024/12
Y1 - 2024/12
N2 - Text-to-image Person Re-Identification (TIReID) is an identity retrieval task between visual and textual modalities. Previous research focuses on learning rich and diverse modality-shared semantic features and achieving excellent performance. However, they still have several notable limitations: (1)Noisy influence: Due to the difficulty of cross-modality annotation and the uncertainty of crowdsourced label quality, it is inevitable to introduce noisy labels by incorrect text–image pairs. (2)Sample imbalance: Datasets collected from real-world sources often face an unbalanced distribution of samples across different categories, which results in inconsistent parameters update progress with the training phrase. To address these issues, we propose a two-stage training pipeline for TIReID learning with noisy correspondence. Firstly, we employ a Noisy Correspondence Detector based on heterogeneous relation retrieval estimating confidence weights from each sample pair. Secondly, we design a multi-view triplet loss function, which leverages sample-level features to interact with global class centers, addressing sample imbalance and facilitating a smoother distribution in feature space. Finally, we utilize these clean samples to train the model through a progressive learning process. Extensive experiments on RSTPReid, CUHK-PEDES, and ICFG-PEDES demonstrate the effectiveness of our method against the state-of-the-art TIReID methods.
AB - Text-to-image Person Re-Identification (TIReID) is an identity retrieval task between visual and textual modalities. Previous research focuses on learning rich and diverse modality-shared semantic features and achieving excellent performance. However, they still have several notable limitations: (1)Noisy influence: Due to the difficulty of cross-modality annotation and the uncertainty of crowdsourced label quality, it is inevitable to introduce noisy labels by incorrect text–image pairs. (2)Sample imbalance: Datasets collected from real-world sources often face an unbalanced distribution of samples across different categories, which results in inconsistent parameters update progress with the training phrase. To address these issues, we propose a two-stage training pipeline for TIReID learning with noisy correspondence. Firstly, we employ a Noisy Correspondence Detector based on heterogeneous relation retrieval estimating confidence weights from each sample pair. Secondly, we design a multi-view triplet loss function, which leverages sample-level features to interact with global class centers, addressing sample imbalance and facilitating a smoother distribution in feature space. Finally, we utilize these clean samples to train the model through a progressive learning process. Extensive experiments on RSTPReid, CUHK-PEDES, and ICFG-PEDES demonstrate the effectiveness of our method against the state-of-the-art TIReID methods.
KW - Cross-modality retrieval
KW - Noisy detection
KW - Text-to-image person re-identification
UR - https://www.scopus.com/pages/publications/85209187884
U2 - 10.1007/s00530-024-01530-6
DO - 10.1007/s00530-024-01530-6
M3 - 文章
AN - SCOPUS:85209187884
SN - 0942-4962
VL - 30
JO - Multimedia Systems
JF - Multimedia Systems
IS - 6
M1 - 338
ER -