跳到主要导航 跳到搜索 跳到主要内容

Spatial-Temporal Separable Attention for Video Action Recognition

  • Xi Guo
  • , Yikun Hu
  • , Fang Chen
  • , Yuhui Jin
  • , Jian Qiao
  • , Jian Huang*
  • , Qin Yang
  • *此作品的通讯作者
  • Beihang University
  • Meituan

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Convolutional neural networks (CNNs) have been proved as a efficient method for various of visual recognition tasks. However, it is more difficult for CNNs to capture long-range spatial-temporal cues in dynamic videos than in static images. Recent nonlocal neural networks attempt to overcome this problem by a self-attention mechanism, where pair-wise affinities for all the spatial-temporal positions are calculated. However, this introduces a substantial computational burden. In this paper, we propose a spatial-temporal separable attention module (STSAM) to reduce the computational complexity. The experimental results, based on the Kinetics 400 benchmark, show that our model achieves better performance but introduces less extra FLOPs than nonlocal neural networks.

源语言英语
主期刊名Proceedings - 2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022
出版商Institute of Electrical and Electronics Engineers Inc.
224-228
页数5
ISBN(电子版)9781665473644
DOI
出版状态已出版 - 2022
活动2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022 - Virtual, Online, 中国
期限: 19 7月 202221 7月 2022

出版系列

姓名Proceedings - 2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022

会议

会议2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022
国家/地区中国
Virtual, Online
时期19/07/2221/07/22

指纹

探究 'Spatial-Temporal Separable Attention for Video Action Recognition' 的科研主题。它们共同构成独一无二的指纹。

引用此