TY - GEN
T1 - MF-BERT
T2 - 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
AU - Shi, Jianxin
AU - Chen, Jinhao
AU - Chen, Xiaolong
AU - Ma, Jun
AU - Wo, Tianyu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Accurately predicting the future motions of traffic agents is essential for autonomous systems. Despite the significant success of existing motion forecasting methods based on supervised learning, they still exhibit two main limitations. First, when annotated data for a scene is limited, these methods often fail to achieve the expected accuracy. Second, they typically rely on complex architectures and extensive prior knowledge to improve performance. To overcome these challenges, we propose MF-BERT, a novel framework that adapts the concept of BERT to motion forecasting, inspired by advancements in the self-supervised pre-training paradigm. During pre-training, we design a siamese sequence modeling task with an asymmetric mask strategy to capture complex behavior patterns of agents. During fine-tuning, the pre-trained representation module initializes the feature encoder of the motion forecasting model, and a multimodal trajectory decoder generates all possible predictions. Experimental results demonstrate the superiority of MF-BERT over state-of-the-art methods.
AB - Accurately predicting the future motions of traffic agents is essential for autonomous systems. Despite the significant success of existing motion forecasting methods based on supervised learning, they still exhibit two main limitations. First, when annotated data for a scene is limited, these methods often fail to achieve the expected accuracy. Second, they typically rely on complex architectures and extensive prior knowledge to improve performance. To overcome these challenges, we propose MF-BERT, a novel framework that adapts the concept of BERT to motion forecasting, inspired by advancements in the self-supervised pre-training paradigm. During pre-training, we design a siamese sequence modeling task with an asymmetric mask strategy to capture complex behavior patterns of agents. During fine-tuning, the pre-trained representation module initializes the feature encoder of the motion forecasting model, and a multimodal trajectory decoder generates all possible predictions. Experimental results demonstrate the superiority of MF-BERT over state-of-the-art methods.
KW - Autonomous Motion Forecasting
KW - Self-supervised Learning
KW - Time Series Analysis
UR - https://www.scopus.com/pages/publications/105003865797
U2 - 10.1109/ICASSP49660.2025.10889386
DO - 10.1109/ICASSP49660.2025.10889386
M3 - 会议稿件
AN - SCOPUS:105003865797
T3 - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
BT - 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Proceedings
A2 - Rao, Bhaskar D
A2 - Trancoso, Isabel
A2 - Sharma, Gaurav
A2 - Mehta, Neelesh B.
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 April 2025 through 11 April 2025
ER -