TY - JOUR
T1 - Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object Detection
AU - Wang, Xuehao
AU - Li, Shuai
AU - Chen, Chenglizhao
AU - Fang, Yuming
AU - Hao, Aimin
AU - Qin, Hong
N1 - Publisher Copyright:
© 1992-2012 IEEE.
PY - 2021
Y1 - 2021
N2 - Existing RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance.
AB - Existing RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance.
KW - RGB-D saliency detection
KW - data-level fusion
KW - lightweight designed triple-stream network
UR - https://www.scopus.com/pages/publications/85096814100
U2 - 10.1109/TIP.2020.3037470
DO - 10.1109/TIP.2020.3037470
M3 - 文章
C2 - 33201813
AN - SCOPUS:85096814100
SN - 1057-7149
VL - 30
SP - 458
EP - 471
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
M1 - 9261990
ER -