TY - JOUR
T1 - SignMask
T2 - Structure-aware Masked Modeling for Holistic 3D Sign Language Production
AU - Xia, Yibo
AU - Zhan, Qihui
AU - Luo, Xiaoyan
AU - Shi, Xiaofeng
AU - Wang, Yunhong
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.
PY - 2026/1/12
Y1 - 2026/1/12
N2 - Sign Language Production (SLP) aims to translate spoken textual languages into sign language sequences, which can significantly bridge the communication gap for deaf and hard-of-hearing individuals. Most previous SLP methods typically rely on skeleton-based data, which hinders their realism and expressive capacity. In this work, we address expressive 3D SLP tasks to generate high-quality 3D holistic sign motions driven by spoken language. However, existing 3D SLP methods struggle to accurately capture spatial relationships within intricate 3D structures and overlook the alignment of semantics at the word level. To overcome these limitations, we propose SignMask, a novel generative masked modeling framework that enhances spatial structure awareness and semantic understanding. We first design a structural holistic sign motion tokenizer that hierarchically learns discrete tokens of body and hand movements. This tokenizer adaptively aggregates 3D SMPL-X pose features corresponding to the same semantic parts and dynamically adjusts the weights between pose features of different semantic parts, enhancing spatial awareness and ensuring semantic consistency. Building on these tokenized representations, we introduce a specialized Sign-M Transformer to learn masked token prediction guided by textual input. Our Sign-M Transformer employs a hierarchical masking strategy, alongside spatio-temporal and cross-modal attention mechanisms, to effectively capture complex spatio-temporal relationships among sign tokens and semantic dependencies between sign and text tokens. During inference, our SignMask model parallelly and iteratively fills up the missing motion tokens starting from full-masked token sequences, therefore achieving high-fidelity and efficient 3D sign avatar generation. Extensive experiments demonstrate the superior performance of our approach compared to existing SLP methods across various lingual sign language datasets in generating high-quality and semantically consistent sign language motions.
AB - Sign Language Production (SLP) aims to translate spoken textual languages into sign language sequences, which can significantly bridge the communication gap for deaf and hard-of-hearing individuals. Most previous SLP methods typically rely on skeleton-based data, which hinders their realism and expressive capacity. In this work, we address expressive 3D SLP tasks to generate high-quality 3D holistic sign motions driven by spoken language. However, existing 3D SLP methods struggle to accurately capture spatial relationships within intricate 3D structures and overlook the alignment of semantics at the word level. To overcome these limitations, we propose SignMask, a novel generative masked modeling framework that enhances spatial structure awareness and semantic understanding. We first design a structural holistic sign motion tokenizer that hierarchically learns discrete tokens of body and hand movements. This tokenizer adaptively aggregates 3D SMPL-X pose features corresponding to the same semantic parts and dynamically adjusts the weights between pose features of different semantic parts, enhancing spatial awareness and ensuring semantic consistency. Building on these tokenized representations, we introduce a specialized Sign-M Transformer to learn masked token prediction guided by textual input. Our Sign-M Transformer employs a hierarchical masking strategy, alongside spatio-temporal and cross-modal attention mechanisms, to effectively capture complex spatio-temporal relationships among sign tokens and semantic dependencies between sign and text tokens. During inference, our SignMask model parallelly and iteratively fills up the missing motion tokens starting from full-masked token sequences, therefore achieving high-fidelity and efficient 3D sign avatar generation. Extensive experiments demonstrate the superior performance of our approach compared to existing SLP methods across various lingual sign language datasets in generating high-quality and semantically consistent sign language motions.
KW - Generative Masked Model
KW - Sign Avatar
KW - Sign Language Production
UR - https://www.scopus.com/pages/publications/105028903053
U2 - 10.1145/3776750
DO - 10.1145/3776750
M3 - 文章
AN - SCOPUS:105028903053
SN - 1551-6857
VL - 22
JO - ACM Transactions on Multimedia Computing, Communications and Applications
JF - ACM Transactions on Multimedia Computing, Communications and Applications
IS - 1
M1 - 22
ER -