TY - JOUR
T1 - LDFE
T2 - Laplacian Decoupled Feature Enhancement block for dual-stream CNN-based RGB-IR object detection
AU - Dong, Wenhao
AU - Luo, Xiaoyan
AU - Yang, Linlin
AU - Zhu, Haodong
AU - Shi, Xiaorong
AU - Guo, Guodong
AU - Zhang, Baochang
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/11
Y1 - 2026/11
N2 - The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN backbones built upon YOLO for feature extraction and focus on the design of feature fusion. In this paper, we introduce the Laplacian Decoupled Feature Enhancement block (LDFE) to fuse features from different stages of the dual-stream CNN backbones. By design, LDFE simultaneously considers the characteristics of modalities and structures for feature fusion by employing global–local decomposition, denoising, fusion, and reconstruction, sequentially. The LDFE first separates features into global and local components based on Laplacian Pyramid, and then performs denoising and fusion based on Global State Space Enhancement module (GS2E) and Local Convolutional Correlation Enhancement module (LC2E) separately. Specifically, the GS2E conducts a two-branch architecture for the main and auxiliary modalities. It dynamically suppresses noise in the main modality through cross-modal attention derived from the auxiliary modality, while employing a State Space Model to capture long-range dependencies within the global feature representations of the main modality. To obtain bidirectional interaction, the two modalities systematically alternate their main/auxiliary roles. Moreover, the LC2E suppresses noise in local features and leverages spatial and channel dimension along with triple convolution to extract fine-grained details for fusion. These innovative designs achieve a significant performance improvement, with mAP surpassing the SOTA methods 6.2%, 3.7%, 4.7%, 2.3%, 4.1% and 2.0% on M3FD, DroneVehicle, LLVIP, FLIR-Aligned, KAIST and VEDAI datasets, respectively.
AB - The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN backbones built upon YOLO for feature extraction and focus on the design of feature fusion. In this paper, we introduce the Laplacian Decoupled Feature Enhancement block (LDFE) to fuse features from different stages of the dual-stream CNN backbones. By design, LDFE simultaneously considers the characteristics of modalities and structures for feature fusion by employing global–local decomposition, denoising, fusion, and reconstruction, sequentially. The LDFE first separates features into global and local components based on Laplacian Pyramid, and then performs denoising and fusion based on Global State Space Enhancement module (GS2E) and Local Convolutional Correlation Enhancement module (LC2E) separately. Specifically, the GS2E conducts a two-branch architecture for the main and auxiliary modalities. It dynamically suppresses noise in the main modality through cross-modal attention derived from the auxiliary modality, while employing a State Space Model to capture long-range dependencies within the global feature representations of the main modality. To obtain bidirectional interaction, the two modalities systematically alternate their main/auxiliary roles. Moreover, the LC2E suppresses noise in local features and leverages spatial and channel dimension along with triple convolution to extract fine-grained details for fusion. These innovative designs achieve a significant performance improvement, with mAP surpassing the SOTA methods 6.2%, 3.7%, 4.7%, 2.3%, 4.1% and 2.0% on M3FD, DroneVehicle, LLVIP, FLIR-Aligned, KAIST and VEDAI datasets, respectively.
KW - Dual-stream backbone
KW - Global and local feature fusion
KW - Laplacian Pyramid
KW - RGB-IR object detection
UR - https://www.scopus.com/pages/publications/105038250917
U2 - 10.1016/j.patcog.2026.113935
DO - 10.1016/j.patcog.2026.113935
M3 - 文章
AN - SCOPUS:105038250917
SN - 0031-3203
VL - 179
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 113935
ER -