Abstract
Due to their unique elevated perspective, roadside sensors can provide extended field-of-view perception information for autonomous vehicles, compensating for the limited sensing range of onboard sensors. However, technical challenges such as the precise detection of small targets under complex weather conditions must be overcome. To address these challenges, a novel model, CDFF-DETR, is proposed based on the Detection Transformer (DETR) architecture, which employs a Cross-scale Dual Feature Fusion(CDFF) pyramid network to enhance small object detection performance. The model incorporates a Lightweight Gated Feature Aggregation(LGFA) unit alongside an Intra-scale Feature Enhancement and Interaction(IFEI) module, significantly reducing computational demands while enhancing small object detection accuracy. Experimental results on the Rope3D dataset demonstrate that compared with the baseline model, CDFF-DETR improves APtiny by 5.1% and APsmall by 4.4% while reducing network parameters by 38.2%. This outperforms representative YOLO and RT-DETR models, which are based on Convolutional Neural Network(CNN) and Vision Transformer(ViT), respectively. Furthermore, the trained model maintains its superiority in small object detection for two other popular datasets, proving its generalizability across training and testing domain.
| Original language | English |
|---|---|
| Article number | 106346 |
| Journal | Digital Signal Processing: A Review Journal |
| Volume | 183 |
| DOIs | |
| State | Published - 1 Nov 2026 |
Keywords
- Dual feature fusion
- Lightweight
- Roadside data
- Small object detection
Fingerprint
Dive into the research topics of 'CDFF-DETR: A lightweight small object detection framework based on cross-scale dual feature fusion pyramid network for roadside data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver