TY - JOUR
T1 - LGFCTR
T2 - Local and Global Feature Convolutional Transformer for Image Matching
AU - Zhong, Wenhao
AU - Jiang, Jie
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2025/4/25
Y1 - 2025/4/25
N2 - Image matching that establishing correspondences across images is a challenging task under extreme conditions. Capturing local and global features simultaneously is an important way to mitigate such an issue according to the human-intuitive insight that human vision relies on both local appearances and global structures for matching. However, recent methods were still stuck in two issues: (1) CNN-based encoders only extract local features; (2) transformers lack locality and rely on explicit positional encoding. Inspired by the locality and implicit positional encoding of convolutions, a novel convolutional transformer LGFCTR is proposed to capture both local contexts and global structures more sufficiently for detector-free matching. Firstly, a universal FPN-like framework captures global structures in self-encoder as well as cross-decoder by transformers and compensates local contexts as well as implicit positional encoding by convolutions. Secondly, a novel Convolutional Transformer Module explores multi-scale long-range dependencies in a grouped manner and further aggregates local information as well as imposing neighborhood constraints for locality enhancement. Finally, a novel regression-based Sub-pixel Refinement Module exploits the whole fine-grained window features for fine-level positional deviation regression. LGFCTR achieves establishing robust and accurate correspondences across illumination variations, viewpoint changes, and scale differences. LGFCTR outperforms state-of-the-art methods by 1.14% and 0.7% in AUC@10°respectively on two general benchmarks of relative pose estimation. LGFCTR also surpasses the second place by 0.0223 on a benchmark of feature matching. LGFCTR consistently demonstrates superior performances on a wide range of benchmarks, including feature matching (HPatches), homography estimation (HPatches), relative pose estimation (MegaDepth and YFCC100M), and visual localization (Aachen Day-Night v1.1).
AB - Image matching that establishing correspondences across images is a challenging task under extreme conditions. Capturing local and global features simultaneously is an important way to mitigate such an issue according to the human-intuitive insight that human vision relies on both local appearances and global structures for matching. However, recent methods were still stuck in two issues: (1) CNN-based encoders only extract local features; (2) transformers lack locality and rely on explicit positional encoding. Inspired by the locality and implicit positional encoding of convolutions, a novel convolutional transformer LGFCTR is proposed to capture both local contexts and global structures more sufficiently for detector-free matching. Firstly, a universal FPN-like framework captures global structures in self-encoder as well as cross-decoder by transformers and compensates local contexts as well as implicit positional encoding by convolutions. Secondly, a novel Convolutional Transformer Module explores multi-scale long-range dependencies in a grouped manner and further aggregates local information as well as imposing neighborhood constraints for locality enhancement. Finally, a novel regression-based Sub-pixel Refinement Module exploits the whole fine-grained window features for fine-level positional deviation regression. LGFCTR achieves establishing robust and accurate correspondences across illumination variations, viewpoint changes, and scale differences. LGFCTR outperforms state-of-the-art methods by 1.14% and 0.7% in AUC@10°respectively on two general benchmarks of relative pose estimation. LGFCTR also surpasses the second place by 0.0223 on a benchmark of feature matching. LGFCTR consistently demonstrates superior performances on a wide range of benchmarks, including feature matching (HPatches), homography estimation (HPatches), relative pose estimation (MegaDepth and YFCC100M), and visual localization (Aachen Day-Night v1.1).
KW - Feature matching
KW - Image matching
KW - Pose estimation
KW - Transformers
UR - https://www.scopus.com/pages/publications/85215397357
U2 - 10.1016/j.eswa.2025.126393
DO - 10.1016/j.eswa.2025.126393
M3 - 文章
AN - SCOPUS:85215397357
SN - 0957-4174
VL - 270
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 126393
ER -