TY - JOUR
T1 - StagNet
T2 - An Attentive Semantic RNN for Group Activity and Individual Action Recognition
AU - Qi, Mengshi
AU - Wang, Yunhong
AU - Qin, Jie
AU - Li, Annan
AU - Luo, Jiebo
AU - Van Gool, Luc
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2020/2
Y1 - 2020/2
N2 - In real life, group activity recognition plays a significant and fundamental role in a variety of applications, e.g. sports video analysis, abnormal behavior detection, and intelligent surveillance. In a complex dynamic scene, a crucial yet challenging issue is how to better model the spatio-temporal contextual information and inter-person relationship. In this paper, we present a novel attentive semantic recurrent neural network (RNN), namely, stagNet, for understanding group activities and individual actions in videos, by combining the spatio-temporal attention mechanism and semantic graph modeling. Specifically, a structured semantic graph is explicitly modeled to express the spatial contextual content of the whole scene, which is further incorporated with the temporal factor through structural-RNN. By virtue of the 'factor sharing' and 'message passing' mechanisms, our stagNet is capable of extracting discriminative and informative spatio-temporal representations and capturing inter-person relationships. Moreover, we adopt a spatio-temporal attention model to focus on key persons/frames for improved recognition performance. Besides, a body-region attention and a global-part feature pooling strategy are devised for individual action recognition. In experiments, four widely-used public datasets are adopted for performance evaluation, and the extensive results demonstrate the superiority and effectiveness of our method.
AB - In real life, group activity recognition plays a significant and fundamental role in a variety of applications, e.g. sports video analysis, abnormal behavior detection, and intelligent surveillance. In a complex dynamic scene, a crucial yet challenging issue is how to better model the spatio-temporal contextual information and inter-person relationship. In this paper, we present a novel attentive semantic recurrent neural network (RNN), namely, stagNet, for understanding group activities and individual actions in videos, by combining the spatio-temporal attention mechanism and semantic graph modeling. Specifically, a structured semantic graph is explicitly modeled to express the spatial contextual content of the whole scene, which is further incorporated with the temporal factor through structural-RNN. By virtue of the 'factor sharing' and 'message passing' mechanisms, our stagNet is capable of extracting discriminative and informative spatio-temporal representations and capturing inter-person relationships. Moreover, we adopt a spatio-temporal attention model to focus on key persons/frames for improved recognition performance. Besides, a body-region attention and a global-part feature pooling strategy are devised for individual action recognition. In experiments, four widely-used public datasets are adopted for performance evaluation, and the extensive results demonstrate the superiority and effectiveness of our method.
KW - Action Recognition
KW - Group Activity Recognition
KW - RNN
KW - Scene Understanding
KW - Semantic Graph
KW - Spatio-temporal Attention
UR - https://www.scopus.com/pages/publications/85075215755
U2 - 10.1109/TCSVT.2019.2894161
DO - 10.1109/TCSVT.2019.2894161
M3 - 文章
AN - SCOPUS:85075215755
SN - 1051-8215
VL - 30
SP - 549
EP - 565
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
IS - 2
M1 - 8621027
ER -