TY - GEN
T1 - Pose-attention
T2 - 2nd International Conference on Computer Vision, Image, and Deep Learning
AU - He, Zhijun
AU - Zhao, Hongbo
AU - Yi, Na
AU - Feng, Wenquan
N1 - Publisher Copyright:
© 2021 SPIE.
PY - 2021
Y1 - 2021
N2 - This paper proposes a novel baseline for deep person ReID methods by introducing human pose-based attention mechanism. Benefiting from deep convolutional network, there has been great progress of person re-identification (ReID) in recent years, which aims at retrieving the same person identities from images captured by different cameras. Most of existing methods focus on designing complex network structures to achieve higher scores on public datasets, but few works pay attention to baseline design. A strong baseline is crucial in experiments and could make the elaborated proposed methods more convincing. The present study makes use of a pre-trained human pose estimator to extract human key-point information. Then, we propose a novel manner to fuse pose information with global feature from Resnet50, which could lead the network concentrate more on discriminative key-point feature areas. Our work could achieve 94.8% rank-1 accuracy & 87.4% mean average precision (mAP) on Market1501, and outperform all other existing baselines that only use Resnet50 to our best knowledge. What's more, experiment results also suggest that with the help of pose information, our work could naturally be robust against misalignment and occlusion problems.
AB - This paper proposes a novel baseline for deep person ReID methods by introducing human pose-based attention mechanism. Benefiting from deep convolutional network, there has been great progress of person re-identification (ReID) in recent years, which aims at retrieving the same person identities from images captured by different cameras. Most of existing methods focus on designing complex network structures to achieve higher scores on public datasets, but few works pay attention to baseline design. A strong baseline is crucial in experiments and could make the elaborated proposed methods more convincing. The present study makes use of a pre-trained human pose estimator to extract human key-point information. Then, we propose a novel manner to fuse pose information with global feature from Resnet50, which could lead the network concentrate more on discriminative key-point feature areas. Our work could achieve 94.8% rank-1 accuracy & 87.4% mean average precision (mAP) on Market1501, and outperform all other existing baselines that only use Resnet50 to our best knowledge. What's more, experiment results also suggest that with the help of pose information, our work could naturally be robust against misalignment and occlusion problems.
KW - Convolutional Network
KW - Deep Learning
KW - Human Pose Estimation
KW - Person Re-identification
KW - Smart Video surveillance
UR - https://www.scopus.com/pages/publications/85118466753
U2 - 10.1117/12.2604694
DO - 10.1117/12.2604694
M3 - 会议稿件
AN - SCOPUS:85118466753
T3 - Proceedings of SPIE - The International Society for Optical Engineering
BT - 2nd International Conference on Computer Vision, Image, and Deep Learning
A2 - bin Ahmad, Badrul Hisham
A2 - Cen, Fengjie
PB - SPIE
Y2 - 25 June 2021 through 27 June 2021
ER -