Skip to main navigation Skip to search Skip to main content

Multi-Scale Spatial-Temporal Representation Learning for mmWave Radar 3D Human Pose Estimation

  • Yaxin Li
  • , Jun Wang
  • , Changshun Yuan*
  • , Xiaoming Yuan
  • , Yuquan Luo
  • , Song Liang
  • *Corresponding author for this work
  • Beihang University
  • Northeastern University China

Research output: Contribution to journalArticlepeer-review

Abstract

Millimeter-wave (mmWave) radar has emerged as a promising sensing modality for 3D human pose estimation in ubiquitous Internet of Things (IoT) applications, owing to its privacy-preserving nature and robustness under challenging illumination conditions. However, accurate skeletal reconstruction remains difficult due to the inherent sparsity, non-uniform distribution, and instability of radar point clouds. To address these challenges, this paper proposes MS-STPoseNet, a unified multi-scale spatio-temporal learning framework for robust mmWavebased human pose estimation. For spatial representation, MS-STPoseNet employs a hierarchical multi-scale spatial encoder built upon PointNet++, which leverages multi-scale grouping to effectively capture body structures across different spatial resolutions, from fine-grained joint regions to limb- and torso-level configurations, under irregular radar observations. For temporal modeling, a multi-branch Temporal Convolutional Network (TCN) with different dilation rates is introduced to model multi-rate motion dynamics, enabling effective representation of both rapid limb movements and smoother torso motions. An attention mechanism is further incorporated to enhance informative temporal features while suppressing noise. The entire framework is trained end-to-end to estimate per-frame 3D poses by exploiting short-term temporal context, thereby improving robustness under noisy sensing conditions. Extensive experiments conducted on a self-collected dataset and two public benchmarks demonstrate that MS-STPoseNet consistently outperforms state-of-the-art methods in terms of pose estimation accuracy and cross-subject generalization, achieving an MPJPE of 3.08 cm on the self-collected dataset. In addition, the proposed framework exhibits favorable computational efficiency and a compact model size, highlighting its potential applicability to practical IoT sensing systems.

Original languageEnglish
JournalIEEE Internet of Things Journal
DOIs
StateAccepted/In press - 2026

Keywords

  • 3D Human Pose Estimation
  • Deep Learning
  • Internet of Things
  • Millimeter-wave Radar
  • Point Cloud
  • Spatio-temporal Learning

Fingerprint

Dive into the research topics of 'Multi-Scale Spatial-Temporal Representation Learning for mmWave Radar 3D Human Pose Estimation'. Together they form a unique fingerprint.

Cite this