TY - JOUR
T1 - ScoreSeg
T2 - Leveraging Score-Based Generative Model for Self-Supervised Semantic Segmentation of Remote Sensing
AU - Lu, Junzhe
AU - He, Guangjun
AU - Dou, Hongkun
AU - Gao, Qing
AU - Fang, Leyuan
AU - Deng, Yue
N1 - Publisher Copyright:
© 2008-2012 IEEE.
PY - 2023
Y1 - 2023
N2 - The performance of semantic segmentation of remote sensing images (RSIs) heavily depends on the number of pixel-level annotations. In practice, the accumulation of pixel-level annotations for large RSIs is quite expensive or even impossible under certain scenarios. Here, we try to solve this data-intensive problem from the novel aspect of score-based self-supervise learning (SSL) and introduce a robust RSI semantic segmentation model called ScoreSeg. Unlike traditional pixel-level SSL paradigms, the generative SSL mechanism in ScoreSeg is simple in loss design and stable in pretraining, granting it an indispensable ability in dense feature learning from very large RSIs. In the model implementation, ScoreSeg first extracts pixelwise representations of RSIs by pretraining a time-dependent score-based model on abundant off-the-shelf unlabeled RSIs. Then, to address the sparse feature problem in RSIs, the collected features from different timesteps and resolutions are aggregated together forming a rich feature map for downstream semantic segmentation. Experimental results on three datasets show that our proposed ScoreSeg outperforms state-of-the-art (SOTA) SSL methods and alternative pretraining models on ImageNet by nontrivial margins, especially with very limited annotations.
AB - The performance of semantic segmentation of remote sensing images (RSIs) heavily depends on the number of pixel-level annotations. In practice, the accumulation of pixel-level annotations for large RSIs is quite expensive or even impossible under certain scenarios. Here, we try to solve this data-intensive problem from the novel aspect of score-based self-supervise learning (SSL) and introduce a robust RSI semantic segmentation model called ScoreSeg. Unlike traditional pixel-level SSL paradigms, the generative SSL mechanism in ScoreSeg is simple in loss design and stable in pretraining, granting it an indispensable ability in dense feature learning from very large RSIs. In the model implementation, ScoreSeg first extracts pixelwise representations of RSIs by pretraining a time-dependent score-based model on abundant off-the-shelf unlabeled RSIs. Then, to address the sparse feature problem in RSIs, the collected features from different timesteps and resolutions are aggregated together forming a rich feature map for downstream semantic segmentation. Experimental results on three datasets show that our proposed ScoreSeg outperforms state-of-the-art (SOTA) SSL methods and alternative pretraining models on ImageNet by nontrivial margins, especially with very limited annotations.
KW - Remote sensing
KW - score-based generative model
KW - self-supervised learning
KW - semantic segmentation
UR - https://www.scopus.com/pages/publications/85171776054
U2 - 10.1109/JSTARS.2023.3314866
DO - 10.1109/JSTARS.2023.3314866
M3 - 文章
AN - SCOPUS:85171776054
SN - 1939-1404
VL - 16
SP - 8818
EP - 8833
JO - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
JF - IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
ER -