摘要
Depth estimation from a single image is critical in robot navigation and scene understanding. It is also a complex problem in computer vision. Aiming at the inaccurate depth estimation of a single image, we propose a single-image depth estimation method based on ViT. First, we downsample the image by the pre-trained DenseNet and encode the features into sequences suitable for ViT. Then, the densely connected ViT processes the global context information, and the feature sequence is reassembled into high-dimensional feature maps. Finally, Upsampling to obtain a complete depth image. We conduct comparative experiments with some reprsentative depth estimation methods on the NYU V2 dataset, and ablation experiments on the network structure. This paper quantitatively analyzes the average relative error, root means square error, and other errors. The results show that the method can generate high-quality depth images with rich details for a single image. Compared with the traditional encoder-decoder method, the PSNR value of the proposed method is increased by 1.052 dB on average, the REL is decreased by 7.7%–21.8%, and the RMS is reduced by 5.6%–16.9%.
| 投稿的翻译标题 | A High-Quality Depth Estimation Method for Single Image |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 1761-1770 |
| 页数 | 10 |
| 期刊 | Jisuanji Fuzhu Sheji Yu Tuxingxue Xuebao/Journal of Computer-Aided Design and Computer Graphics |
| 卷 | 36 |
| 期 | 11 |
| DOI | |
| 出版状态 | 已出版 - 2024 |
关键词
- attention mechanism
- deep learning
- depth estimation
- vision Transformer
学术指纹
探究 '面向单幅图像的高质量深度估计方法' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver