跳到主要导航 跳到搜索 跳到主要内容

LGFCTR: Local and Global Feature Convolutional Transformer for Image Matching

  • Wenhao Zhong
  • , Jie Jiang*
  • *此作品的通讯作者
  • Beihang University

科研成果: 期刊稿件文章同行评审

摘要

Image matching that establishing correspondences across images is a challenging task under extreme conditions. Capturing local and global features simultaneously is an important way to mitigate such an issue according to the human-intuitive insight that human vision relies on both local appearances and global structures for matching. However, recent methods were still stuck in two issues: (1) CNN-based encoders only extract local features; (2) transformers lack locality and rely on explicit positional encoding. Inspired by the locality and implicit positional encoding of convolutions, a novel convolutional transformer LGFCTR is proposed to capture both local contexts and global structures more sufficiently for detector-free matching. Firstly, a universal FPN-like framework captures global structures in self-encoder as well as cross-decoder by transformers and compensates local contexts as well as implicit positional encoding by convolutions. Secondly, a novel Convolutional Transformer Module explores multi-scale long-range dependencies in a grouped manner and further aggregates local information as well as imposing neighborhood constraints for locality enhancement. Finally, a novel regression-based Sub-pixel Refinement Module exploits the whole fine-grained window features for fine-level positional deviation regression. LGFCTR achieves establishing robust and accurate correspondences across illumination variations, viewpoint changes, and scale differences. LGFCTR outperforms state-of-the-art methods by 1.14% and 0.7% in AUC@10°respectively on two general benchmarks of relative pose estimation. LGFCTR also surpasses the second place by 0.0223 on a benchmark of feature matching. LGFCTR consistently demonstrates superior performances on a wide range of benchmarks, including feature matching (HPatches), homography estimation (HPatches), relative pose estimation (MegaDepth and YFCC100M), and visual localization (Aachen Day-Night v1.1).

源语言英语
文章编号126393
期刊Expert Systems with Applications
270
DOI
出版状态已出版 - 25 4月 2025

指纹

探究 'LGFCTR: Local and Global Feature Convolutional Transformer for Image Matching' 的科研主题。它们共同构成独一无二的指纹。

引用此