Abstract
Multimodal knowledge graph completion (MKGC) focuses on predicting missing head or tail entities in multimodal knowledge graphs. While previous studies have emphasized the integration of images and descriptive sentences to enhance entity representation learning, our research highlights that not all multimodal information is beneficial for MKGC. In fact, irrelevant multimodal data can lead to incorrect predictions and diminished performance. This finding underscores the need for a balance between leveraging relevant multimodal information and minimizing the impact of irrelevant data in knowledge graph representation learning. To address this challenge, we propose a self-adaptive fusion model for MKGC (SAFKGC). SAFKGC utilizes a cross-modal transformer with an innovative input sequence to select features and assess the importance of different modalities. It also identifies irrelevant multimodal information based on prediction confidence, weakening its influence while maintaining reasoning consistency between multimodal and unimodal models through soft-label learning. Experimental results across four widely used MKGC datasets demonstrate that our model achieves competitive or superior performance compared to entity-aware the state-of-the-art approaches. Additional experiments reveal that SAFKGC can dynamically adjust unimodal KGC results to enhance multimodal KGC outcomes. The code is available at https://anonymous.4open.science/r/SAFKGC-BBE4/.
| Original language | English |
|---|---|
| Article number | 132003 |
| Journal | Expert Systems with Applications |
| Volume | 322 |
| DOIs | |
| State | Published - 1 Aug 2026 |
Keywords
- Knowledge graph completion
- Multimodal learning
Fingerprint
Dive into the research topics of 'Revealing multimodal information trade-off in multimodal knowledge graph completion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver