TY - JOUR
T1 - Mask-Guided Feature Routing and Adaptive Context Modeling for Wide-FoV UAV Object Detection in IoT Remote Sensing
AU - Wu, Lingfan
AU - Feng, Yachun
AU - Zhang, Hong
AU - Li, Yawei
N1 - Publisher Copyright:
© 2026 by the authors.
PY - 2026/6
Y1 - 2026/6
N2 - Object detection in wide-field-of-view (wide-FoV) unmanned aerial vehicle (UAV) imagery for Internet of Things (IoT) remote sensing applications requires accurate recognition of tiny objects under severe background redundancy and extreme scale variation. As the field of view expands, conventional dense detectors tend to waste substantial computation on non-informative regions, while feature downsampling and static receptive fields often cause the dilution of foreground information and scale confusion. To address these issues, we propose MFRC-Det, a unified framework built upon two complementary principles: mask-guided feature routing and adaptive context modeling. Specifically, a Superpixel-Masking Generator (SP-Masker) is introduced to estimate an image-space soft foreground prior by comparing Simple Linear Iterative Clustering (SLIC) superpixel histograms with a peripheral background reference, propagating the resulting scores on a superpixel adjacency graph, and projecting the refined region-level scores back to a pixel-level routing mask. Guided by these priors, a Greedy-Cutter (G-Cutter) converts dense feature maps into compact, foreground-focused patches without repeated backbone evaluation on cropped image regions, thereby reducing redundant background computation while preserving local structural coherence. On top of the retained regions, an Adaptive Receptive-field Selection Network (ARSNet) aggregates multi-scale contextual responses from several learnable receptive-field candidate branches. ARSNet predicts spatial selection weights conditioned on the input features, allowing each location to emphasize a suitable receptive-field response for object representation. Experimental results on VisDrone-DET and UAVDT demonstrate that MFRC-Det achieves competitive detection accuracy with favorable computational efficiency. Specifically, MFRC-Det obtains 36.1% AP, 60.4% (Formula presented.), and 38.5 FPS on VisDrone-DET and 21.3% AP, 36.8% (Formula presented.), and 37.4 FPS on UAVDT. These results validate the effectiveness of mask-guided feature routing and adaptive context modeling for wide-FoV UAV object detection and suggest their potential value for computation-efficient aerial perception in IoT remote sensing applications.
AB - Object detection in wide-field-of-view (wide-FoV) unmanned aerial vehicle (UAV) imagery for Internet of Things (IoT) remote sensing applications requires accurate recognition of tiny objects under severe background redundancy and extreme scale variation. As the field of view expands, conventional dense detectors tend to waste substantial computation on non-informative regions, while feature downsampling and static receptive fields often cause the dilution of foreground information and scale confusion. To address these issues, we propose MFRC-Det, a unified framework built upon two complementary principles: mask-guided feature routing and adaptive context modeling. Specifically, a Superpixel-Masking Generator (SP-Masker) is introduced to estimate an image-space soft foreground prior by comparing Simple Linear Iterative Clustering (SLIC) superpixel histograms with a peripheral background reference, propagating the resulting scores on a superpixel adjacency graph, and projecting the refined region-level scores back to a pixel-level routing mask. Guided by these priors, a Greedy-Cutter (G-Cutter) converts dense feature maps into compact, foreground-focused patches without repeated backbone evaluation on cropped image regions, thereby reducing redundant background computation while preserving local structural coherence. On top of the retained regions, an Adaptive Receptive-field Selection Network (ARSNet) aggregates multi-scale contextual responses from several learnable receptive-field candidate branches. ARSNet predicts spatial selection weights conditioned on the input features, allowing each location to emphasize a suitable receptive-field response for object representation. Experimental results on VisDrone-DET and UAVDT demonstrate that MFRC-Det achieves competitive detection accuracy with favorable computational efficiency. Specifically, MFRC-Det obtains 36.1% AP, 60.4% (Formula presented.), and 38.5 FPS on VisDrone-DET and 21.3% AP, 36.8% (Formula presented.), and 37.4 FPS on UAVDT. These results validate the effectiveness of mask-guided feature routing and adaptive context modeling for wide-FoV UAV object detection and suggest their potential value for computation-efficient aerial perception in IoT remote sensing applications.
KW - IoT remote sensing
KW - adaptive context modeling
KW - mask-guided feature routing
KW - tiny object detection
KW - unmanned aerial vehicle
KW - wide-FoV object detection
UR - https://www.scopus.com/pages/publications/105041460083
U2 - 10.3390/rs18111753
DO - 10.3390/rs18111753
M3 - 文章
AN - SCOPUS:105041460083
SN - 2072-4292
VL - 18
JO - Remote Sensing
JF - Remote Sensing
IS - 11
M1 - 1753
ER -