TY - JOUR
T1 - Modeling latent cross-modal interaction for reliability-aware entity alignment in multimodal relation extraction
AU - Liu, Xiyang
AU - Hu, Chunming
AU - Zhang, Richong
AU - Chen, Junfan
AU - Xu, Baowen
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/9/5
Y1 - 2026/9/5
N2 - Multimodal relation extraction aims to exploit the complementary contextual information across modalities to identify semantic relationships between entities. While prior works have achieved promising results by applying various alignment strategies to enhance cross-modal fusion, their effectiveness remains constrained by the absence of entity-centric alignment annotations. Recently developed region-level multimodal large language models (MLLMs) show potential as automated data annotators capable of providing diverse cross-modal alignment labels. However, MLLMs can produce unreliable grounding outputs, especially when processing an unfamiliar entity in ambiguous context. To address this limitation, we propose a latent cross-modal information interaction modeling framework to acquire entity alignment annotations with quantified reliability scores. Our approach introduces token-wise weight backtracking equipped with logit-guided causal attention aggregation to learn input–output dependencies. Sequence-level evidence detection subsequently analyzes parametric alignment cues to derive validated cross-modal entity mappings. Lightweight task model is adaptively optimized with learned alignment supervision to enhance its ability to bridge modality gaps. Empirical studies show that our reliability-aware alignment annotations significantly improve performance on multimodal relation extraction tasks. The source code is available at https://github.com/liuxiyang641/LCIM.
AB - Multimodal relation extraction aims to exploit the complementary contextual information across modalities to identify semantic relationships between entities. While prior works have achieved promising results by applying various alignment strategies to enhance cross-modal fusion, their effectiveness remains constrained by the absence of entity-centric alignment annotations. Recently developed region-level multimodal large language models (MLLMs) show potential as automated data annotators capable of providing diverse cross-modal alignment labels. However, MLLMs can produce unreliable grounding outputs, especially when processing an unfamiliar entity in ambiguous context. To address this limitation, we propose a latent cross-modal information interaction modeling framework to acquire entity alignment annotations with quantified reliability scores. Our approach introduces token-wise weight backtracking equipped with logit-guided causal attention aggregation to learn input–output dependencies. Sequence-level evidence detection subsequently analyzes parametric alignment cues to derive validated cross-modal entity mappings. Lightweight task model is adaptively optimized with learned alignment supervision to enhance its ability to bridge modality gaps. Empirical studies show that our reliability-aware alignment annotations significantly improve performance on multimodal relation extraction tasks. The source code is available at https://github.com/liuxiyang641/LCIM.
KW - Cross-modal entity alignment
KW - Information extraction
KW - Multimodal large language model
KW - Multimodal relation extraction
UR - https://www.scopus.com/pages/publications/105042670952
U2 - 10.1016/j.knosys.2026.116474
DO - 10.1016/j.knosys.2026.116474
M3 - 文章
AN - SCOPUS:105042670952
SN - 0950-7051
VL - 349
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116474
ER -