TY - GEN
T1 - Automatic 4D facial expression recognition using dynamic geometrical image network
AU - Li, Weijian
AU - Huang, Di
AU - Li, Huibin
AU - Wang, Yunhong
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/6/5
Y1 - 2018/6/5
N2 - In this paper, we propose a novel Dynamic Geometrical Image Network (DGIN) for automatic 4D Facial Expression Recognition (FER). Given a 3D video represented as a sequence of face scans, we first estimate their differential geometry quantities and generate geometrical images, including Depth Images (DPI), three Normal Component Images (NCI) and Shape Index Images (SII). These geometrical images are then fed into DGIN for end-to-end training and prediction. DGIN consists of a short-term temporal pooling layer for dynamic geometric image generation, several repetitions of convolution+ReLU+pooling layers for facial spatial feature extraction, and a long-term temporal pooling layer for dynamic feature map fusion, followed by fully connected layers and a joint loss layer. During the training phase, the two-stage longterm and short-term sliding window scheme is introduced for data augmentation and temporal pooling. Meanwhile, a joint loss integrating both the cross-entropy loss and the triplet loss is used to achieve more discriminative expression features. In the testing phase, only the short-term sliding window scheme is applied to the whole video sequence of certain geometric images, whose outputs further go through the deep net for expression similarity measurement. The final result is achieved by fusing the predicted expression scores of different types of geometrical images. Experimental results reported on the BU- 4DFE database demonstrate the effectiveness of the proposed approach.
AB - In this paper, we propose a novel Dynamic Geometrical Image Network (DGIN) for automatic 4D Facial Expression Recognition (FER). Given a 3D video represented as a sequence of face scans, we first estimate their differential geometry quantities and generate geometrical images, including Depth Images (DPI), three Normal Component Images (NCI) and Shape Index Images (SII). These geometrical images are then fed into DGIN for end-to-end training and prediction. DGIN consists of a short-term temporal pooling layer for dynamic geometric image generation, several repetitions of convolution+ReLU+pooling layers for facial spatial feature extraction, and a long-term temporal pooling layer for dynamic feature map fusion, followed by fully connected layers and a joint loss layer. During the training phase, the two-stage longterm and short-term sliding window scheme is introduced for data augmentation and temporal pooling. Meanwhile, a joint loss integrating both the cross-entropy loss and the triplet loss is used to achieve more discriminative expression features. In the testing phase, only the short-term sliding window scheme is applied to the whole video sequence of certain geometric images, whose outputs further go through the deep net for expression similarity measurement. The final result is achieved by fusing the predicted expression scores of different types of geometrical images. Experimental results reported on the BU- 4DFE database demonstrate the effectiveness of the proposed approach.
KW - 3D facial expression recognition
KW - Data augmentation.
KW - Deep neural network
KW - Joint-loss-function
UR - https://www.scopus.com/pages/publications/85049393771
U2 - 10.1109/FG.2018.00014
DO - 10.1109/FG.2018.00014
M3 - 会议稿件
AN - SCOPUS:85049393771
T3 - Proceedings - 13th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2018
SP - 24
EP - 30
BT - Proceedings - 13th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2018
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 13th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2018
Y2 - 15 May 2018 through 19 May 2018
ER -