TY - GEN
T1 - Multi-modal Trajectory Prediction Network that Integrates Historical Motion and Spatio-Temporal Interaction
AU - Li, Chenlong
AU - Li, Mingxing
AU - Zhao, Jian
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Multi-modal trajectory prediction (MTP) has become the research trend in the field of autonomous driving, as it provides multiple plausible trajectories. However, related works lack attention to the temporal dependence of historical features and inherent association between multiple trajectory modes, which may lead to large deviations. To address these critical limitations, we propose a multi-modal trajectory prediction network that integrates historical motion and spatio-temporal interaction (MTPN-IMI). In MTPN-IMI, a local spatio-temporal graph (LSTG) is constructed to model local agent-agent interaction. Furthermore, a Causal Convolution Module (CCM) and a Causal Self-Attention Module (CSAM) are introduced to focus on local and global temporal dependence in historical motion feature and local agent-agent interaction feature. Moreover, a Cross Attention Module (CAM) is utilized to capture the inherent association between multiple modes. Experiments show that our model outperforms related models on Argoverse1.1 validation set, achieving superior prediction accuracy.
AB - Multi-modal trajectory prediction (MTP) has become the research trend in the field of autonomous driving, as it provides multiple plausible trajectories. However, related works lack attention to the temporal dependence of historical features and inherent association between multiple trajectory modes, which may lead to large deviations. To address these critical limitations, we propose a multi-modal trajectory prediction network that integrates historical motion and spatio-temporal interaction (MTPN-IMI). In MTPN-IMI, a local spatio-temporal graph (LSTG) is constructed to model local agent-agent interaction. Furthermore, a Causal Convolution Module (CCM) and a Causal Self-Attention Module (CSAM) are introduced to focus on local and global temporal dependence in historical motion feature and local agent-agent interaction feature. Moreover, a Cross Attention Module (CAM) is utilized to capture the inherent association between multiple modes. Experiments show that our model outperforms related models on Argoverse1.1 validation set, achieving superior prediction accuracy.
KW - Multi-modal
KW - Spatio-temporal interaction
KW - Temporal dependence
KW - Trajectory prediction
UR - https://www.scopus.com/pages/publications/105038951421
U2 - 10.1007/978-981-95-6553-5_39
DO - 10.1007/978-981-95-6553-5_39
M3 - 会议稿件
AN - SCOPUS:105038951421
SN - 9789819565528
T3 - Lecture Notes in Electrical Engineering
SP - 433
EP - 441
BT - Proceedings of 2025 Chinese Intelligent Systems Conference
A2 - Jia, Yingmin
A2 - Liu, Yang
A2 - Zhang, Weicun
A2 - Fu, Yongling
PB - Springer Science and Business Media Deutschland GmbH
T2 - 21st Chinese Intelligent Systems Conference, CISC 2025
Y2 - 25 October 2025 through 26 October 2025
ER -