TY - JOUR
T1 - MCCANet
T2 - A Precision and Efficient Bottom Tracking Method Based on Cross-Cue Fusion of Single and Multiple Ping Inputs
AU - Zhang, Yanxian
AU - Huo, Guanying
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2025
Y1 - 2025
N2 - The primary purpose of bottom tracking is to identify the boundary between the water column area and the image area in the side scan sonar (SSS) waterfall map. However, noise in the water column, often caused by complex measurement environments, poses significant challenges for automatic bottom tracking. Therefore, we propose a multihead cross-cue attention network (MCCANet), a lightweight network designed to achieve precision and efficient bottom tracking. MCCANet consists of four modules: the input module, encoder, feature fusion module, and decoder. The input module extracts features from one ping and five consecutive pings while maintaining dimensionally consistent outputs. The encoder employs simple 1-D convolutional layers to extract features from 1-D sequences. The feature fusion module fuses and enhances features from single and multiple pings using multihead cross-cue attention (MCCA) mechanism. Finally, the decoder reconstructs the dimensionality and maps the inputs to semantic labels. To train and evaluate the model, we annotate the NY_HudsonRiver_sss-xtf open-source dataset. Compared to the best-performing single-ping bottom tracking method, MCCANet achieves significant improvements in both intersection over union (IoU) and Dice metrics. It reduces the mean offset error (MOE) and total root-mean-square error (TRMSE) by 45.49% and 30.71%, respectively, while achieving a prediction speed of 1985 p/s. Additionally, MCCANet demonstrates robust performance on SSS datasets collected from Yunnan and Taiwan province, further validating its generalization capability. Crucially, experimental results confirm that MCCANet exhibits high robustness to noise.
AB - The primary purpose of bottom tracking is to identify the boundary between the water column area and the image area in the side scan sonar (SSS) waterfall map. However, noise in the water column, often caused by complex measurement environments, poses significant challenges for automatic bottom tracking. Therefore, we propose a multihead cross-cue attention network (MCCANet), a lightweight network designed to achieve precision and efficient bottom tracking. MCCANet consists of four modules: the input module, encoder, feature fusion module, and decoder. The input module extracts features from one ping and five consecutive pings while maintaining dimensionally consistent outputs. The encoder employs simple 1-D convolutional layers to extract features from 1-D sequences. The feature fusion module fuses and enhances features from single and multiple pings using multihead cross-cue attention (MCCA) mechanism. Finally, the decoder reconstructs the dimensionality and maps the inputs to semantic labels. To train and evaluate the model, we annotate the NY_HudsonRiver_sss-xtf open-source dataset. Compared to the best-performing single-ping bottom tracking method, MCCANet achieves significant improvements in both intersection over union (IoU) and Dice metrics. It reduces the mean offset error (MOE) and total root-mean-square error (TRMSE) by 45.49% and 30.71%, respectively, while achieving a prediction speed of 1985 p/s. Additionally, MCCANet demonstrates robust performance on SSS datasets collected from Yunnan and Taiwan province, further validating its generalization capability. Crucially, experimental results confirm that MCCANet exhibits high robustness to noise.
KW - 1-D sequence
KW - bottom tracking
KW - feature fusion
KW - multihead cross-cue attention network (MCCANet)
KW - side scan sonar (SSS)
UR - https://www.scopus.com/pages/publications/105003032448
U2 - 10.1109/TGRS.2025.3553565
DO - 10.1109/TGRS.2025.3553565
M3 - 文章
AN - SCOPUS:105003032448
SN - 0196-2892
VL - 63
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5909215
ER -