Skip to main navigation Skip to search Skip to main content

Spatial-Temporal Separable Attention for Video Action Recognition

  • Xi Guo
  • , Yikun Hu
  • , Fang Chen
  • , Yuhui Jin
  • , Jian Qiao
  • , Jian Huang*
  • , Qin Yang
  • *Corresponding author for this work
  • Beihang University
  • Meituan

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Convolutional neural networks (CNNs) have been proved as a efficient method for various of visual recognition tasks. However, it is more difficult for CNNs to capture long-range spatial-temporal cues in dynamic videos than in static images. Recent nonlocal neural networks attempt to overcome this problem by a self-attention mechanism, where pair-wise affinities for all the spatial-temporal positions are calculated. However, this introduces a substantial computational burden. In this paper, we propose a spatial-temporal separable attention module (STSAM) to reduce the computational complexity. The experimental results, based on the Kinetics 400 benchmark, show that our model achieves better performance but introduces less extra FLOPs than nonlocal neural networks.

Original languageEnglish
Title of host publicationProceedings - 2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages224-228
Number of pages5
ISBN (Electronic)9781665473644
DOIs
StatePublished - 2022
Event2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022 - Virtual, Online, China
Duration: 19 Jul 202221 Jul 2022

Publication series

NameProceedings - 2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022

Conference

Conference2022 International Conference on Frontiers of Artificial Intelligence and Machine Learning, FAIML 2022
Country/TerritoryChina
CityVirtual, Online
Period19/07/2221/07/22

Keywords

  • attention mechanism
  • convolutional neural network
  • nonlocal neural networks
  • video action recognition

Fingerprint

Dive into the research topics of 'Spatial-Temporal Separable Attention for Video Action Recognition'. Together they form a unique fingerprint.

Cite this