跳到主要导航 跳到搜索 跳到主要内容

SIM-MultiDepth: Self-Supervised Indoor Monocular Multi-Frame Depth Estimation Based on Texture-Aware Masking

  • Beihang University
  • Nanchang Institute of Technology

科研成果: 期刊稿件文章同行评审

摘要

Self-supervised monocular depth estimation methods have become the focus of research since ground truth data are not required. Current single-image-based works only leverage appearance-based features, thus achieving a limited performance. Deep learning based multiview stereo works facilitate the research on multi-frame depth estimation methods. Some multi-frame methods build cost volumes and take multiple frames as inputs at the time of test to fully utilize geometric cues between adjacent frames. Nevertheless, low-textured regions, which are dominant in indoor scenes, tend to cause unreliable depth hypotheses in the cost volume. Few self-supervised multi-frame methods have been used to conduct research on the issue of low-texture areas in indoor scenes. To handle this issue, we propose SIM-MultiDepth, a self-supervised indoor monocular multi-frame depth estimation framework. A self-supervised single-frame depth estimation network is introduced to learn the relative poses and supervise the multi-frame depth learning. A texture-aware depth consistency loss is designed considering the calculation of the patch-based photometric loss. Only the areas where multi-frame depth prediction is considered unreliable in low-texture regions are supervised by the single-frame network. This approach helps improve the depth estimation accuracy. The experimental results on the NYU Depth V2 dataset validate the effectiveness of SIM-MultiDepth. The zero-shot generalization studies on the 7-Scenes and Campus Indoor datasets aid in the analysis of the application characteristics of SIM-MultiDepth.

源语言英语
文章编号2221
期刊Remote Sensing
16
12
DOI
出版状态已出版 - 6月 2024

学术指纹

探究 'SIM-MultiDepth: Self-Supervised Indoor Monocular Multi-Frame Depth Estimation Based on Texture-Aware Masking' 的科研主题。它们共同构成独一无二的学术指纹。

引用此