Skip to main navigation Skip to search Skip to main content

Modeling latent cross-modal interaction for reliability-aware entity alignment in multimodal relation extraction

  • Beihang University
  • Ltd.

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal relation extraction aims to exploit the complementary contextual information across modalities to identify semantic relationships between entities. While prior works have achieved promising results by applying various alignment strategies to enhance cross-modal fusion, their effectiveness remains constrained by the absence of entity-centric alignment annotations. Recently developed region-level multimodal large language models (MLLMs) show potential as automated data annotators capable of providing diverse cross-modal alignment labels. However, MLLMs can produce unreliable grounding outputs, especially when processing an unfamiliar entity in ambiguous context. To address this limitation, we propose a latent cross-modal information interaction modeling framework to acquire entity alignment annotations with quantified reliability scores. Our approach introduces token-wise weight backtracking equipped with logit-guided causal attention aggregation to learn input–output dependencies. Sequence-level evidence detection subsequently analyzes parametric alignment cues to derive validated cross-modal entity mappings. Lightweight task model is adaptively optimized with learned alignment supervision to enhance its ability to bridge modality gaps. Empirical studies show that our reliability-aware alignment annotations significantly improve performance on multimodal relation extraction tasks. The source code is available at https://github.com/liuxiyang641/LCIM.

Original languageEnglish
Article number116474
JournalKnowledge-Based Systems
Volume349
DOIs
StatePublished - 5 Sep 2026

Keywords

  • Cross-modal entity alignment
  • Information extraction
  • Multimodal large language model
  • Multimodal relation extraction

Fingerprint

Dive into the research topics of 'Modeling latent cross-modal interaction for reliability-aware entity alignment in multimodal relation extraction'. Together they form a unique fingerprint.

Cite this