Skip to main navigation Skip to search Skip to main content

Multimodal Object Detection by Adaptive Channel Enhancement and Attention Fusion

  • Yaqi Mei*
  • , Tianyuan Zhang
  • , Huobin Tan*
  • *Corresponding author for this work
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Cross-modal feature fusion is a critical research area in multimodal object detection, focusing on integrating features extracted from different modalities to retain richer semantic information. While several advanced fusion strategies have been proposed, most fail to effectively address the interaction of complementary information between modalities, resulting in suboptimal information exchange and fusion that do not fully leverage the intrinsic characteristics of each modality. To address these challenges, this paper introduces an Adaptive Channel Enhancement and Attention Fusion (ACAF) module, which bridges the feature gaps across modalities, enabling smooth information interaction and attention-based multimodal feature integration. Specifically, the module adaptively reweights weaker channels in each modality using features from other modalities, thereby enhancing the expressiveness of single-modal feature representations. Additionally, an attention mechanism is employed to capture multidimensional cross-modal relationships, facilitating efficient feature fusion. Experimental results on the DroneVehicle and VEDAI datasets demonstrate that our method significantly outperforms the baseline models, achieving improvements of 3.0% and 2.4% in recall, 2.1% and 1.7% in mAP@50, and 3.0% and 4.6% in mAP@50:95, respectively. This shows that channel reweighting enhances cross-modal information interaction and fusion, leading to superior performance in multimodal object detection.

Original languageEnglish
Title of host publicationInternational Joint Conference on Neural Networks, IJCNN 2025 - Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331510428
DOIs
StatePublished - 2025
Event2025 International Joint Conference on Neural Networks, IJCNN 2025 - Rome, Italy
Duration: 30 Jun 20255 Jul 2025

Publication series

NameProceedings of the International Joint Conference on Neural Networks
ISSN (Print)2161-4393
ISSN (Electronic)2161-4407

Conference

Conference2025 International Joint Conference on Neural Networks, IJCNN 2025
Country/TerritoryItaly
CityRome
Period30/06/255/07/25

Keywords

  • channel enhancement
  • cross-modal complementarity
  • feature fusion
  • Multimodal object detection

Fingerprint

Dive into the research topics of 'Multimodal Object Detection by Adaptive Channel Enhancement and Attention Fusion'. Together they form a unique fingerprint.

Cite this