TY - GEN
T1 - Multimodal Object Detection by Adaptive Channel Enhancement and Attention Fusion
AU - Mei, Yaqi
AU - Zhang, Tianyuan
AU - Tan, Huobin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Cross-modal feature fusion is a critical research area in multimodal object detection, focusing on integrating features extracted from different modalities to retain richer semantic information. While several advanced fusion strategies have been proposed, most fail to effectively address the interaction of complementary information between modalities, resulting in suboptimal information exchange and fusion that do not fully leverage the intrinsic characteristics of each modality. To address these challenges, this paper introduces an Adaptive Channel Enhancement and Attention Fusion (ACAF) module, which bridges the feature gaps across modalities, enabling smooth information interaction and attention-based multimodal feature integration. Specifically, the module adaptively reweights weaker channels in each modality using features from other modalities, thereby enhancing the expressiveness of single-modal feature representations. Additionally, an attention mechanism is employed to capture multidimensional cross-modal relationships, facilitating efficient feature fusion. Experimental results on the DroneVehicle and VEDAI datasets demonstrate that our method significantly outperforms the baseline models, achieving improvements of 3.0% and 2.4% in recall, 2.1% and 1.7% in mAP@50, and 3.0% and 4.6% in mAP@50:95, respectively. This shows that channel reweighting enhances cross-modal information interaction and fusion, leading to superior performance in multimodal object detection.
AB - Cross-modal feature fusion is a critical research area in multimodal object detection, focusing on integrating features extracted from different modalities to retain richer semantic information. While several advanced fusion strategies have been proposed, most fail to effectively address the interaction of complementary information between modalities, resulting in suboptimal information exchange and fusion that do not fully leverage the intrinsic characteristics of each modality. To address these challenges, this paper introduces an Adaptive Channel Enhancement and Attention Fusion (ACAF) module, which bridges the feature gaps across modalities, enabling smooth information interaction and attention-based multimodal feature integration. Specifically, the module adaptively reweights weaker channels in each modality using features from other modalities, thereby enhancing the expressiveness of single-modal feature representations. Additionally, an attention mechanism is employed to capture multidimensional cross-modal relationships, facilitating efficient feature fusion. Experimental results on the DroneVehicle and VEDAI datasets demonstrate that our method significantly outperforms the baseline models, achieving improvements of 3.0% and 2.4% in recall, 2.1% and 1.7% in mAP@50, and 3.0% and 4.6% in mAP@50:95, respectively. This shows that channel reweighting enhances cross-modal information interaction and fusion, leading to superior performance in multimodal object detection.
KW - channel enhancement
KW - cross-modal complementarity
KW - feature fusion
KW - Multimodal object detection
UR - https://www.scopus.com/pages/publications/105023969939
U2 - 10.1109/IJCNN64981.2025.11227959
DO - 10.1109/IJCNN64981.2025.11227959
M3 - 会议稿件
AN - SCOPUS:105023969939
T3 - Proceedings of the International Joint Conference on Neural Networks
BT - International Joint Conference on Neural Networks, IJCNN 2025 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 International Joint Conference on Neural Networks, IJCNN 2025
Y2 - 30 June 2025 through 5 July 2025
ER -