TY - JOUR
T1 - Learning cross-modal correlations by exploring inter-word semantics and stacked co-attention
AU - Yu, Jing
AU - Lu, Yuhang
AU - Zhang, Weifeng
AU - Qin, Zengchang
AU - Liu, Yanbing
AU - Hu, Yue
N1 - Publisher Copyright:
© 2018 Elsevier B.V.
PY - 2020/2
Y1 - 2020/2
N2 - Cross-modal information retrieval aims to find heterogeneous data of various modalities from a given query of one modality. The main challenge is to learn the semantic correlations between different modalities and measure the distance across modalities. For text-image retrieval, existing work mostly uses off-the-shelf Convolutional Neural Network (CNN) for image feature extraction. For texts, word-level features such as bag-of-words or word2vec are employed to build deep learning models to represent texts. Besides word-level semantics, the semantic relations between words are also informative but less explored. In this paper, we explore the inter-word semantics by modelling texts by graphs using similarity measure based on word2vec. Besides feature presentations, we further study the problem of information imbalance between different modalities when describing the same semantics. For example textual descriptions often contain more background information that cannot be conveyed by images and vice versa. We propose a stacked co-attention network to progressively learn the mutually attended features of different modalities and enhance their fine-grained correlations. A dual-path neural network is proposed for cross-modal information retrieval. The model is trained by a pairwise similarity loss function to maximize the similarity of relevant text-image pairs and minimize the similarity of irrelevant pairs. Experimental results show that the proposed model outperforms the state-of-the-art methods significantly, with 19% improvement on accuracy for the best case.
AB - Cross-modal information retrieval aims to find heterogeneous data of various modalities from a given query of one modality. The main challenge is to learn the semantic correlations between different modalities and measure the distance across modalities. For text-image retrieval, existing work mostly uses off-the-shelf Convolutional Neural Network (CNN) for image feature extraction. For texts, word-level features such as bag-of-words or word2vec are employed to build deep learning models to represent texts. Besides word-level semantics, the semantic relations between words are also informative but less explored. In this paper, we explore the inter-word semantics by modelling texts by graphs using similarity measure based on word2vec. Besides feature presentations, we further study the problem of information imbalance between different modalities when describing the same semantics. For example textual descriptions often contain more background information that cannot be conveyed by images and vice versa. We propose a stacked co-attention network to progressively learn the mutually attended features of different modalities and enhance their fine-grained correlations. A dual-path neural network is proposed for cross-modal information retrieval. The model is trained by a pairwise similarity loss function to maximize the similarity of relevant text-image pairs and minimize the similarity of irrelevant pairs. Experimental results show that the proposed model outperforms the state-of-the-art methods significantly, with 19% improvement on accuracy for the best case.
KW - Cross-modal correlation
KW - Cross-modal retrieval
KW - Fine-grained correlation
KW - Inter-word semantics
KW - Stacked co-attention
UR - https://www.scopus.com/pages/publications/85054390142
U2 - 10.1016/j.patrec.2018.08.017
DO - 10.1016/j.patrec.2018.08.017
M3 - 文章
AN - SCOPUS:85054390142
SN - 0167-8655
VL - 130
SP - 189
EP - 198
JO - Pattern Recognition Letters
JF - Pattern Recognition Letters
ER -