TY - JOUR
T1 - Additional encoder is not what you need
T2 - Unleashing the potential of CLIP for weakly-supervised semantic segmentation
AU - Lv, You
AU - Kang, Guoliang
AU - Wei, Wei
N1 - Publisher Copyright:
© 2025 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/4/15
Y1 - 2026/4/15
N2 - In this paper, we focus on weakly supervised semantic segmentation (WSSS) task, which aims to generate accurate pixel-level predictions with only image-level labels available. A typical solution is to generate dense masks from sparse attention maps with the guidance of image-level labels. Recent works utilize Contrastive Language-Image Pre-training (CLIP) in WSSS task. As CLIP aligns visual features with textural features, most previous works directly adopt CLIP to generate class-aware attention maps. However, they usually train a separate segmentation network to obtain satisfactory performance with attention maps as pseudo masks. In this work, we find that a frozen CLIP as segmentation encoder may perform quite well in WSSS task, and additional encoder training is not necessary. Specifically, we first generate pixel-level pseudo masks via Grad-CAM from the frozen CLIP model. To obtain fine-grained attention, we compute CAMs by jointly leveraging gradients from both the last and intermediate transformer layers. Subsequently, we append a lightweight convolutional decoder to produce initial segmentation predictions. To mitigate the class-preference and space-preference biases inherent in CLIP, the Bias Rectification Module is employed to adaptively rectify these biases. We further impose a contrastive loss between masked image features and textual features to better align the visual and textual representations within CLIP. In addition, a supervised loss between the generated pseudo masks and the rectified predictions is employed to refine the dense outputs. Extensive experiments on PASCAL VOC2012 and MS COCO2014 datasets demonstrate that our method performs favorably against previous state-of-the-art WSSS approaches.
AB - In this paper, we focus on weakly supervised semantic segmentation (WSSS) task, which aims to generate accurate pixel-level predictions with only image-level labels available. A typical solution is to generate dense masks from sparse attention maps with the guidance of image-level labels. Recent works utilize Contrastive Language-Image Pre-training (CLIP) in WSSS task. As CLIP aligns visual features with textural features, most previous works directly adopt CLIP to generate class-aware attention maps. However, they usually train a separate segmentation network to obtain satisfactory performance with attention maps as pseudo masks. In this work, we find that a frozen CLIP as segmentation encoder may perform quite well in WSSS task, and additional encoder training is not necessary. Specifically, we first generate pixel-level pseudo masks via Grad-CAM from the frozen CLIP model. To obtain fine-grained attention, we compute CAMs by jointly leveraging gradients from both the last and intermediate transformer layers. Subsequently, we append a lightweight convolutional decoder to produce initial segmentation predictions. To mitigate the class-preference and space-preference biases inherent in CLIP, the Bias Rectification Module is employed to adaptively rectify these biases. We further impose a contrastive loss between masked image features and textual features to better align the visual and textual representations within CLIP. In addition, a supervised loss between the generated pseudo masks and the rectified predictions is employed to refine the dense outputs. Extensive experiments on PASCAL VOC2012 and MS COCO2014 datasets demonstrate that our method performs favorably against previous state-of-the-art WSSS approaches.
KW - CLIP
KW - Semantic segmentation
KW - Weakly supervised
UR - https://www.scopus.com/pages/publications/105029595406
U2 - 10.1016/j.eswa.2025.130932
DO - 10.1016/j.eswa.2025.130932
M3 - 文章
AN - SCOPUS:105029595406
SN - 0957-4174
VL - 306
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 130932
ER -