TY - GEN
T1 - Improved Sliding Window Smoothing for Video Temporal Action Segmentation and Recognition
AU - Li, Ce
AU - Tian, Yihan
AU - Sheng, Longshuai
AU - Chen, Junzhi
AU - Wang, Tian
AU - Wei, Xianlong
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Despite substantial research on human action segmentation in videos to determine the type and timing of activities, the topic is still unresolved because of the dearth of large-scale annotation data in video analysis applications. Supervised video action segmentation employs a number of temporal convolutional network (TCN) models to address this problem. The process is still difficult, because of the intricate temporal duration division of the movements in the videos. In order to create a soft and flexible video partition, we incorporate an improved sliding window smoothing (ISWS) technique into a TCN baseline model in this study. When screening the target video segmentation sequence, our research method carefully selects three discriminative frames and cleverly integrates them into the adaptive sliding window to more specifically optimize the smoothing effect of the entire prediction sequence. It is particularly worth noting that we implement a doubling penalty mechanism when the window slides to the wrong category position. In order to learn the resultants of effective and ineffective segmentation paths, we designed a new loss function to smooth the candidate frames of the segmentation points in the sliding window using the ISWS scheme. So that our method can increase the receptive field of video segmentation effectively to gain the optimal action segmentation. Experiments on the breakfast, 50salads, and GTEA datasets demonstrate that our method significantly improved the frame accuracy for action segmentation in videos when compared to the state-of-the-art techniques.
AB - Despite substantial research on human action segmentation in videos to determine the type and timing of activities, the topic is still unresolved because of the dearth of large-scale annotation data in video analysis applications. Supervised video action segmentation employs a number of temporal convolutional network (TCN) models to address this problem. The process is still difficult, because of the intricate temporal duration division of the movements in the videos. In order to create a soft and flexible video partition, we incorporate an improved sliding window smoothing (ISWS) technique into a TCN baseline model in this study. When screening the target video segmentation sequence, our research method carefully selects three discriminative frames and cleverly integrates them into the adaptive sliding window to more specifically optimize the smoothing effect of the entire prediction sequence. It is particularly worth noting that we implement a doubling penalty mechanism when the window slides to the wrong category position. In order to learn the resultants of effective and ineffective segmentation paths, we designed a new loss function to smooth the candidate frames of the segmentation points in the sliding window using the ISWS scheme. So that our method can increase the receptive field of video segmentation effectively to gain the optimal action segmentation. Experiments on the breakfast, 50salads, and GTEA datasets demonstrate that our method significantly improved the frame accuracy for action segmentation in videos when compared to the state-of-the-art techniques.
KW - Video segmentation
KW - improved sliding window smoothing
KW - temporal convolutional networks
UR - https://www.scopus.com/pages/publications/85189318597
U2 - 10.1109/CAC59555.2023.10450614
DO - 10.1109/CAC59555.2023.10450614
M3 - 会议稿件
AN - SCOPUS:85189318597
T3 - Proceedings - 2023 China Automation Congress, CAC 2023
SP - 8653
EP - 8658
BT - Proceedings - 2023 China Automation Congress, CAC 2023
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 China Automation Congress, CAC 2023
Y2 - 17 November 2023 through 19 November 2023
ER -