TY - GEN
T1 - YOLP
T2 - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
AU - Fan, Xudong
AU - Zhao, Wei
AU - Zhang, Rufei
AU - Li, Nannan
AU - Li, Dongjin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - While object detection from unmanned aerial vehicles (UAVs) is vital for numerous applications, it confronts two major challenges: detecting small targets with limited effective pixels requires substantial computational resources, and significant feature extraction capacity is wasted on nontarget background regions because the objects are usually sparsely distributed and locally clustered. To improve the detector's robustness and computational efficiency, we propose CFFM-YOLO, an enhanced YOLOv8 detector incorporating a lightweight Cross-scale Feature Fusion Module (CFFM) for dynamic feature fusion across hierarchical feature levels. This module uses dynamic convolutions, scale attention, and channel attention to refine features across spatial, scale, and channel dimensions, boosting efficiency without compromising detection accuracy. For detecting small, non-uniformly distributed objects in complex image backgrounds, methods that employ extra networks for cropping image local regions suffer from high overhead, while increasing input resolution is wasteful due to extensive processing of irrelevant background regions. Building upon CFFM-YOLO with high-resolution inputs, we propose YOLP, which introduces a Feature Patch Slicing Module (FPSM) for selective feature-level cropping rather than processing the entire high-resolution image. The FPSM leverages backbone features to predict an object heatmap, generates size-fixed feature patches, and performs subsequent object detection. FPSM avoids redundant feature extraction caused by extra networks and skips redundant computation on background regions. On the VisDrone and UAVDT benchmarks, the proposed YOLP demonstrates superior detection performance, achieving state-of-the-art m A P scores of 3 8. 2% and 2 5. 5%, respectively. This result substantiates its effectiveness and generalization capability in UAV remote sensing object detection.
AB - While object detection from unmanned aerial vehicles (UAVs) is vital for numerous applications, it confronts two major challenges: detecting small targets with limited effective pixels requires substantial computational resources, and significant feature extraction capacity is wasted on nontarget background regions because the objects are usually sparsely distributed and locally clustered. To improve the detector's robustness and computational efficiency, we propose CFFM-YOLO, an enhanced YOLOv8 detector incorporating a lightweight Cross-scale Feature Fusion Module (CFFM) for dynamic feature fusion across hierarchical feature levels. This module uses dynamic convolutions, scale attention, and channel attention to refine features across spatial, scale, and channel dimensions, boosting efficiency without compromising detection accuracy. For detecting small, non-uniformly distributed objects in complex image backgrounds, methods that employ extra networks for cropping image local regions suffer from high overhead, while increasing input resolution is wasteful due to extensive processing of irrelevant background regions. Building upon CFFM-YOLO with high-resolution inputs, we propose YOLP, which introduces a Feature Patch Slicing Module (FPSM) for selective feature-level cropping rather than processing the entire high-resolution image. The FPSM leverages backbone features to predict an object heatmap, generates size-fixed feature patches, and performs subsequent object detection. FPSM avoids redundant feature extraction caused by extra networks and skips redundant computation on background regions. On the VisDrone and UAVDT benchmarks, the proposed YOLP demonstrates superior detection performance, achieving state-of-the-art m A P scores of 3 8. 2% and 2 5. 5%, respectively. This result substantiates its effectiveness and generalization capability in UAV remote sensing object detection.
KW - Feature fusion
KW - Heatmap prediction
KW - Small object detection
KW - UAV
KW - YOLOv8
UR - https://www.scopus.com/pages/publications/105036004309
U2 - 10.1109/RAAI67517.2025.11423114
DO - 10.1109/RAAI67517.2025.11423114
M3 - 会议稿件
AN - SCOPUS:105036004309
T3 - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
SP - 186
EP - 195
BT - 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence, RAAI 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 December 2025 through 20 December 2025
ER -