TY - JOUR
T1 - Physics-informed visual-inertial mamba for robust train localization in harsh conditions
AU - Xian, Xiaoyu
AU - Zhou, Qiuyang
AU - Tian, Yin
AU - Tian, Daxin
AU - Zhou, Jianshan
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/10
Y1 - 2026/10
N2 - Visual localization is a critical method for achieving reliable train localization, especially in Global Navigation Satellite System (GNSS)-denied environments. However, its effectiveness is severely hampered in harsh operational conditions, such as extreme weather and low-light tunnels, which induce significant visual degradation. To address this fundamental challenge, we propose PI-VIMamba, a novel physics-informed visual-inertial mamba framework. The innovation is centered around a tripartite technical system that integrates a backbone network, a detection module, and a depth estimation module. Firstly, an adaptive channel convolution module is proposed to dramatically enhance feature representation under visual degradation, leading to superior object detection accuracy and more precise depth estimation. Secondly, a kinematics-constrained loss function is introduced leveraging Inertial Measurement Unit (IMU) data. This physical constraint effectively suppresses monocular visual estimation drift, thereby ensuring high-precision, long-distance localization capability even in visually degraded environments. Finally, an improved mamba block is designed for refined general feature extraction, which substantially reduces the network's parameters and computational complexity while accelerating inference. Evaluated on real-world subway line data, PI-VIMamba demonstrably excels in harsh conditions, achieving state-of-the-art performance across key metrics, including detection accuracy, depth estimation precision, and computational efficiency, significantly outperforming existing mainstream models.
AB - Visual localization is a critical method for achieving reliable train localization, especially in Global Navigation Satellite System (GNSS)-denied environments. However, its effectiveness is severely hampered in harsh operational conditions, such as extreme weather and low-light tunnels, which induce significant visual degradation. To address this fundamental challenge, we propose PI-VIMamba, a novel physics-informed visual-inertial mamba framework. The innovation is centered around a tripartite technical system that integrates a backbone network, a detection module, and a depth estimation module. Firstly, an adaptive channel convolution module is proposed to dramatically enhance feature representation under visual degradation, leading to superior object detection accuracy and more precise depth estimation. Secondly, a kinematics-constrained loss function is introduced leveraging Inertial Measurement Unit (IMU) data. This physical constraint effectively suppresses monocular visual estimation drift, thereby ensuring high-precision, long-distance localization capability even in visually degraded environments. Finally, an improved mamba block is designed for refined general feature extraction, which substantially reduces the network's parameters and computational complexity while accelerating inference. Evaluated on real-world subway line data, PI-VIMamba demonstrably excels in harsh conditions, achieving state-of-the-art performance across key metrics, including detection accuracy, depth estimation precision, and computational efficiency, significantly outperforming existing mainstream models.
KW - Harsh conditions
KW - Inertial measurement unit
KW - Monocular vision
KW - Object detection
KW - Physics-informed neural network
KW - Train localization
UR - https://www.scopus.com/pages/publications/105033584622
U2 - 10.1016/j.patcog.2026.113422
DO - 10.1016/j.patcog.2026.113422
M3 - 文章
AN - SCOPUS:105033584622
SN - 0031-3203
VL - 178
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 113422
ER -