TY - GEN
T1 - ScNet
T2 - 2025 IEEE International Conference on Multimedia and Expo, ICME 2025
AU - Shi, Jianxin
AU - Chen, Xiaolong
AU - Xie, Yusen
AU - Chen, Jinhao
AU - Wang, Fali
AU - Ma, Jun
AU - Wo, Tianyu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Predicting the motion of traffic agents is a fundamental challenge in autonomous driving, essential for safe and efficient ego-vehicle planning. Traditional methods typically focus on marginal forecasting, where the trajectory of each agent is predicted separately, leading to inconsistencies in scene-level predictions. To address this issue, we propose a scene-consistency network, named ScNet, which jointly predicts the trajectories of multiple agents in a single feedforward pass, ensuring consistency across all predictions. Our method leverages dual independently initialized student models that interact through cross-network contrastive learning at the global feature level, enhancing robustness and scene consistency in the learned representations. To further improve scene coherence, we incorporate a scene-guided strategy that refines these representations. Additionally, we employ a lightweight, anchor-free decoder that generates predictions for all agents, aligning the forecasts with real-world dynamics. Experiments show significant improvements in multi-world prediction metrics across complex environments. Code and models will be publicly available.
AB - Predicting the motion of traffic agents is a fundamental challenge in autonomous driving, essential for safe and efficient ego-vehicle planning. Traditional methods typically focus on marginal forecasting, where the trajectory of each agent is predicted separately, leading to inconsistencies in scene-level predictions. To address this issue, we propose a scene-consistency network, named ScNet, which jointly predicts the trajectories of multiple agents in a single feedforward pass, ensuring consistency across all predictions. Our method leverages dual independently initialized student models that interact through cross-network contrastive learning at the global feature level, enhancing robustness and scene consistency in the learned representations. To further improve scene coherence, we incorporate a scene-guided strategy that refines these representations. Additionally, we employ a lightweight, anchor-free decoder that generates predictions for all agents, aligning the forecasts with real-world dynamics. Experiments show significant improvements in multi-world prediction metrics across complex environments. Code and models will be publicly available.
KW - Autonomous Driving
KW - Contrastive Representation Learning
KW - Joint Motion Forecasting
KW - Mutual Learning
UR - https://www.scopus.com/pages/publications/105022635132
U2 - 10.1109/ICME59968.2025.11210215
DO - 10.1109/ICME59968.2025.11210215
M3 - 会议稿件
AN - SCOPUS:105022635132
T3 - Proceedings - IEEE International Conference on Multimedia and Expo
BT - 2025 IEEE International Conference on Multimedia and Expo
PB - IEEE Computer Society
Y2 - 30 June 2025 through 4 July 2025
ER -