Abstract
In recent years, deep learning-based stereo matching techniques have achieved significant progress in remote sensing applications such as large-scale scene 3D reconstruction. However, most existing mainstream methods rely on 3D convolutional neural networks to interpret and aggregate the cost volume, and their inherent local receptive fields limit the performance of the model when dealing with complex scenes commonly found in high-resolution satellite imagery, such as weakly-textured and repeated-textured areas. To address this limitation, we propose a novel, plug-and-play Mamba Pyramid Aggregator (MPA). Unlike traditional stacked 3D convolutional structures, MPA combines a pyramid structure with the sequence modeling capability of the Mamba model. While maintaining linear computational complexity, it enhances the network's interpretation ability through global context awareness and fine-grained aggregation at different granularities. MPA can be embedded as a plug-and-play module into existing stereo matching networks. We conducted enhancement experiments on five mainstream models using the WHU-Stereo and US3D satellite remote sensing datasets. The results show that MPA significantly improves the disparity estimation accuracy of all baseline models, particularly excelling in repeated-textured and weakly-textured areas, demonstrating its potential and value in handling complex remote sensing scenarios.
| Original language | English |
|---|---|
| Journal | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Cost Volume Aggregation
- Mamba Model
- Satellite Remote Sensing Images
- Stereo Matching
Fingerprint
Dive into the research topics of 'MPA: A Mamba Pyramid Aggregator for Stereo Matching of Satellite Remote Sensing Images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver