Abstract
Sign Language Production (SLP) aims to translate spoken textual languages into sign language sequences, which can significantly bridge the communication gap for deaf and hard-of-hearing individuals. Most previous SLP methods typically rely on skeleton-based data, which hinders their realism and expressive capacity. In this work, we address expressive 3D SLP tasks to generate high-quality 3D holistic sign motions driven by spoken language. However, existing 3D SLP methods struggle to accurately capture spatial relationships within intricate 3D structures and overlook the alignment of semantics at the word level. To overcome these limitations, we propose SignMask, a novel generative masked modeling framework that enhances spatial structure awareness and semantic understanding. We first design a structural holistic sign motion tokenizer that hierarchically learns discrete tokens of body and hand movements. This tokenizer adaptively aggregates 3D SMPL-X pose features corresponding to the same semantic parts and dynamically adjusts the weights between pose features of different semantic parts, enhancing spatial awareness and ensuring semantic consistency. Building on these tokenized representations, we introduce a specialized Sign-M Transformer to learn masked token prediction guided by textual input. Our Sign-M Transformer employs a hierarchical masking strategy, alongside spatio-temporal and cross-modal attention mechanisms, to effectively capture complex spatio-temporal relationships among sign tokens and semantic dependencies between sign and text tokens. During inference, our SignMask model parallelly and iteratively fills up the missing motion tokens starting from full-masked token sequences, therefore achieving high-fidelity and efficient 3D sign avatar generation. Extensive experiments demonstrate the superior performance of our approach compared to existing SLP methods across various lingual sign language datasets in generating high-quality and semantically consistent sign language motions.
| Original language | English |
|---|---|
| Article number | 22 |
| Journal | ACM Transactions on Multimedia Computing, Communications and Applications |
| Volume | 22 |
| Issue number | 1 |
| DOIs | |
| State | Published - 12 Jan 2026 |
Keywords
- Generative Masked Model
- Sign Avatar
- Sign Language Production
Fingerprint
Dive into the research topics of 'SignMask: Structure-aware Masked Modeling for Holistic 3D Sign Language Production'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver