Skip to main navigation Skip to search Skip to main content

Spatial-temporal Transformer for Skeleton-based Action Recognition

  • Qipeng Zhang
  • , Kexin Liu
  • , Tian Wang*
  • , Peng Shi
  • , Mengyi Zhang
  • , Hichem Snoussi
  • *Corresponding author for this work
  • Beihang University
  • Nanjing University of Science and Technology
  • Fujian Normal University
  • Nanjing Tech University
  • Université de technologie de Troyes

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In the area of skeleton-based human action recognition, GCN has achieved good results in previous research due to its excellent modeling ability on graph data. Recently, transformers have achieved extraordinary results in many computer vision fields. Comparing transformer and GCN, from a certain point of view, we can regard transformer as a kind of dynamic GCN, and the weight of each node is dynamically determined by data. In this work, a three-dimensional position encoding was proposed by us to solve the representation of node spatial information, in order to apply the transformer to the graph data. In addition, similar to Spatial-Temporal Graph Convolutional Networks (ST-GCN), we proposed a Space-Time Transformer (ST-TR), which applies transformers in space and time to extract spatiotemporal feature of skeleton data to complete action recognition.

Original languageEnglish
Title of host publicationProceeding - 2021 China Automation Congress, CAC 2021
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7029-7034
Number of pages6
ISBN (Electronic)9781665426473
DOIs
StatePublished - 2021
Event2021 China Automation Congress, CAC 2021 - Beijing, China
Duration: 22 Oct 202124 Oct 2021

Publication series

NameProceeding - 2021 China Automation Congress, CAC 2021

Conference

Conference2021 China Automation Congress, CAC 2021
Country/TerritoryChina
CityBeijing
Period22/10/2124/10/21

Keywords

  • Action Recognition
  • skeleton data
  • transformers

Fingerprint

Dive into the research topics of 'Spatial-temporal Transformer for Skeleton-based Action Recognition'. Together they form a unique fingerprint.

Cite this