TY - GEN
T1 - Adaptive Feature Aggregation for Video Object Detection
AU - Qian, Yijun
AU - Yu, Lijun
AU - Liu, Wenhe
AU - Kang, Guoliang
AU - Hauptmann, Alexander G.
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/3
Y1 - 2020/3
N2 - Object detection, as a fundamental research topic of computer vision, is facing the challenges of video-related tasks. Objects in videos tend to be blurred, occluded, or out of focus more frequently. Existing works adopt feature aggregation and enhancement to design video-based object detectors. However, most of them do not consider the diversity of object movements and the quality of aggregated context features. Thus, they can not generate comparable results given blurred or crowded videos. In this paper, we propose an adaptive feature aggregation method for video object detection to deal with these problems. We introduce an adaptive quality-similarity weight, with a sparse and dense temporal aggregation policy, into our model. Compared with both image-based and video-based baselines on Im-ageNet and VIRAT datasets, our work consistently demonstrates better performance. Especially, our model improves the average precision of person detection in VIRAT from 85.93% to 87.21%. Several demonstration videos of this work are available.
AB - Object detection, as a fundamental research topic of computer vision, is facing the challenges of video-related tasks. Objects in videos tend to be blurred, occluded, or out of focus more frequently. Existing works adopt feature aggregation and enhancement to design video-based object detectors. However, most of them do not consider the diversity of object movements and the quality of aggregated context features. Thus, they can not generate comparable results given blurred or crowded videos. In this paper, we propose an adaptive feature aggregation method for video object detection to deal with these problems. We introduce an adaptive quality-similarity weight, with a sparse and dense temporal aggregation policy, into our model. Compared with both image-based and video-based baselines on Im-ageNet and VIRAT datasets, our work consistently demonstrates better performance. Especially, our model improves the average precision of person detection in VIRAT from 85.93% to 87.21%. Several demonstration videos of this work are available.
UR - https://www.scopus.com/pages/publications/85085912689
U2 - 10.1109/WACVW50321.2020.9096948
DO - 10.1109/WACVW50321.2020.9096948
M3 - 会议稿件
AN - SCOPUS:85085912689
T3 - Proceedings - 2020 IEEE Winter Conference on Applications of Computer Vision Workshops, WACVW 2020
SP - 143
EP - 147
BT - Proceedings - 2020 IEEE Winter Conference on Applications of Computer Vision Workshops, WACVW 2020
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2020 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2020
Y2 - 1 March 2020 through 5 March 2020
ER -