跳到主要导航 跳到搜索 跳到主要内容

IntSTR: An integrated spatio-temporal relation transformer for video object detection

  • Wentao Zheng*
  • , Hong Zheng
  • , Yuquan Sun
  • , Ying Jing
  • *此作品的通讯作者
  • Beihang University
  • Tsinghua University

科研成果: 期刊稿件文章同行评审

摘要

In recent years, Transformer-based video object detection (VOD) methods have achieved remarkable progress by replacing the hand-crafted components traditionally used in CNN-based detectors. However, many existing approaches rely on staged spatio-temporal modeling strategies, which increase model complexity and restrict early interaction between spatial and temporal information. To overcome these limitations, we propose IntSTR, a novel framework for unified spatio-temporal modeling. At its core, the spatio-temporal relation encoder (STRE) integrates spatio-temporal feature processing within a single encoder through cascaded attention modules. To strengthen temporal consistency, the temporal query relation (TQR) module explicitly captures geometric relations between object queries across adjacent frames with minimal computational overhead. In addition, the Temporal Feature Memory (TFM) maintains a dynamic memory bank that caches temporal contexts, enabling effective feature aggregation and efficient online processing. Extensive experiments on the ImageNet VID dataset validate the effectiveness of our approach. IntSTR achieves an excellent trade-off between accuracy and efficiency, reaching a competitive 87.2 % mAP50 with the ResNet-101 backbone while maintaining real-time performance at 33.4 FPS.

源语言英语
文章编号131704
期刊Neurocomputing
658
DOI
出版状态已出版 - 28 12月 2025

指纹

探究 'IntSTR: An integrated spatio-temporal relation transformer for video object detection' 的科研主题。它们共同构成独一无二的指纹。

引用此