TY - JOUR
T1 - MMT-NET
T2 - a lightweight multi-modal fusion network for UAV target detection in adverse environments
AU - Wang, Chuanyun
AU - Zhou, Mingqi
AU - Sun, Dongdong
AU - Gao, Qian
AU - Li, Zhaokui
AU - Wang, Tian
N1 - Publisher Copyright:
© The Author(s), under exclusive licence to Springer-Verlag GmbH Germany, part of Springer Nature 2025.
PY - 2025/8
Y1 - 2025/8
N2 - To address the challenge of insufficient target detection accuracy for UAVs operating in adverse conditions, such as low illumination, dense fog, and extreme weather, this paper proposes a lightweight multi-modal fusion detection network named MMT-NET, designed to enhance UAV perception capabilities in challenging environments. Built upon the RT-DETR framework, the proposed method employs a dual-branch MobileNetV4 backbone to independently extract features from infrared and visible images. Moreover, a lightweight multi-modal feature interaction module is designed to strengthen the interaction between different modalities, and a lightweight cross-modal attention fusion module is designed to efficiently fuse cross-modal features via a spatial attention mechanism with minimal computational overhead. Extensive experiments on the public multi-modal dataset M3FD demonstrate that MMT-NET achieves 89.9% mAP@50 and 60.3% mAP@50:95, validating its effectiveness in multi-modal detection tasks while maintaining a lightweight architecture. Furthermore, qualitative evaluations under diverse real-world and simulated scenarios—including nighttime, fog, snow, and occlusion—confirm the robustness and generalization capability of the proposed method in complex environments. The source code of this work will be publicly available at: https://github.com/UAVSwarm/MMT-NET.
AB - To address the challenge of insufficient target detection accuracy for UAVs operating in adverse conditions, such as low illumination, dense fog, and extreme weather, this paper proposes a lightweight multi-modal fusion detection network named MMT-NET, designed to enhance UAV perception capabilities in challenging environments. Built upon the RT-DETR framework, the proposed method employs a dual-branch MobileNetV4 backbone to independently extract features from infrared and visible images. Moreover, a lightweight multi-modal feature interaction module is designed to strengthen the interaction between different modalities, and a lightweight cross-modal attention fusion module is designed to efficiently fuse cross-modal features via a spatial attention mechanism with minimal computational overhead. Extensive experiments on the public multi-modal dataset M3FD demonstrate that MMT-NET achieves 89.9% mAP@50 and 60.3% mAP@50:95, validating its effectiveness in multi-modal detection tasks while maintaining a lightweight architecture. Furthermore, qualitative evaluations under diverse real-world and simulated scenarios—including nighttime, fog, snow, and occlusion—confirm the robustness and generalization capability of the proposed method in complex environments. The source code of this work will be publicly available at: https://github.com/UAVSwarm/MMT-NET.
KW - Images attention mechanism
KW - Infrared and visible
KW - Lightweight network
KW - Multi-modal fusion
KW - UAV target detection
UR - https://www.scopus.com/pages/publications/105012481595
U2 - 10.1007/s11554-025-01741-8
DO - 10.1007/s11554-025-01741-8
M3 - 文章
AN - SCOPUS:105012481595
SN - 1861-8200
VL - 22
JO - Journal of Real-Time Image Processing
JF - Journal of Real-Time Image Processing
IS - 4
M1 - 160
ER -