Skip to main navigation Skip to search Skip to main content

MFR-Net: Motion-Guided Feature Refinement Network for Video Small Object Detection in Vision-Based Airport Surveillance Systems

  • Shengjie Zhang
  • , Yang Yang
  • , Yanbo Zhu
  • , Shengsheng Qian
  • , Xiaoxiao Zhang
  • , Kaiquan Cai*
  • *Corresponding author for this work
  • Beihang University
  • State Key Laboratory of CNS/ATM
  • CAS - Institute of Automation
  • University of Chinese Academy of Sciences
  • Aviation Data Communication Corporation

Research output: Contribution to journalArticlepeer-review

Abstract

Video-based surveillance in critical airport surface environments is essential for comprehensive situational awareness, yet it presents a unique and challenging task which we formally define as Video Small Object Detection in Vision-based Airport Surveillance Systems (VSOD-VASS). This task is characterized by a confluence of compounding difficulties, including 1) diminutive object size and severe long-tail imbalance of semantic categories, and 2) consistent modeling of long-range motions and stochastic transient dynamics. To address these challenges, we propose a two-part framework: an offline data augmentation pipeline, named Trajectory-anchored Object Augmentation (TOA), and a Motion-guided Feature Refinement Network (MFR-Net). The TOA pipeline combines trajectory-constrained object generation for controlled appearance diversification and delayed trajectory replay for reusing valid historical object states. The MFR-Net then integrates two key innovations: an adaptive Motion-aware Feature Alignment (MFA) module that robustly models transient dynamics from adjacent frames, and a selective Distant Proximal Temporal Feature Aggregation (DPTFA) module that uses attention to filter noise and aggregate valuable long-range dependencies. To facilitate robust evaluation, we introduce VASSO, a large-scale real-world airport surveillance dataset featuring 1,332,223 instances across 12 categories, where over 87.51% of all objects are small or tiny. Extensive experiments demonstrate that MFR-Net achieves performance with 72.50% mAP and 93.43% mAP_{50}, significantly improving 4.70% mAP with existing SOTA methods. Our work provides a formal problem definition, a strong SOTA method, and a challenging large-scale dataset to advance research in this critical domain.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
StateAccepted/In press - 2026

Keywords

  • Vision-based airport surveillance systems
  • data augmentation
  • temporal feature aggregation
  • transient dynamic modeling
  • video small object detection

Fingerprint

Dive into the research topics of 'MFR-Net: Motion-Guided Feature Refinement Network for Video Small Object Detection in Vision-Based Airport Surveillance Systems'. Together they form a unique fingerprint.

Cite this