跳到主要导航 跳到搜索 跳到主要内容

MMT-NET: a lightweight multi-modal fusion network for UAV target detection in adverse environments

  • Chuanyun Wang*
  • , Mingqi Zhou
  • , Dongdong Sun
  • , Qian Gao
  • , Zhaokui Li
  • , Tian Wang
  • *此作品的通讯作者
  • Shenyang Aerospace University

科研成果: 期刊稿件文章同行评审

摘要

To address the challenge of insufficient target detection accuracy for UAVs operating in adverse conditions, such as low illumination, dense fog, and extreme weather, this paper proposes a lightweight multi-modal fusion detection network named MMT-NET, designed to enhance UAV perception capabilities in challenging environments. Built upon the RT-DETR framework, the proposed method employs a dual-branch MobileNetV4 backbone to independently extract features from infrared and visible images. Moreover, a lightweight multi-modal feature interaction module is designed to strengthen the interaction between different modalities, and a lightweight cross-modal attention fusion module is designed to efficiently fuse cross-modal features via a spatial attention mechanism with minimal computational overhead. Extensive experiments on the public multi-modal dataset M3FD demonstrate that MMT-NET achieves 89.9% mAP@50 and 60.3% mAP@50:95, validating its effectiveness in multi-modal detection tasks while maintaining a lightweight architecture. Furthermore, qualitative evaluations under diverse real-world and simulated scenarios—including nighttime, fog, snow, and occlusion—confirm the robustness and generalization capability of the proposed method in complex environments. The source code of this work will be publicly available at: https://github.com/UAVSwarm/MMT-NET.

源语言英语
文章编号160
期刊Journal of Real-Time Image Processing
22
4
DOI
出版状态已出版 - 8月 2025

学术指纹

探究 'MMT-NET: a lightweight multi-modal fusion network for UAV target detection in adverse environments' 的科研主题。它们共同构成独一无二的学术指纹。

引用此