TY - JOUR
T1 - Exploring Weakly Labeled Images for Video Object Segmentation with Submodular Proposal Selection
AU - Zhang, Yu
AU - Chen, Xiaowu
AU - Li, Jia
AU - Teng, Wei
AU - Song, Haokun
N1 - Publisher Copyright:
© 1992-2012 IEEE.
PY - 2018/9
Y1 - 2018/9
N2 - Video object segmentation (VOS) is important for various computer vision problems, and handling it with minimal human supervision is highly desired for the large-scale applications. To bring down the supervision, existing approaches largely follow a data mining perspective by assuming the availability of multiple videos sharing the same object categories. It, however, would be problematic for the tasks that consume a single video. To address this problem, this paper proposes a novel approach that explores weakly labeled images to solve video object segmentation. Given a video labeled with a target category, images labeled with the same category are collected, from which noisy object exemplars are automatically discovered. After that the proposed approach extracts a set of region proposals on various frames and efficiently matches them with massive noisy exemplars in terms of appearance and spatial context. We then jointly select the best proposals across the video by solving a novel submodular problem that combines region voting and global region matching. Finally, the localization results are leveraged as strong supervision to guide pixel-level segmentation. Extensive experiments are conducted on two challenging public databases: Youtube-Objects and DAVIS. The results suggest that the proposed approach improves over previous weakly supervised/unsupervised approaches significantly, showing a performance even comparable with the several approaches supervised by the costly manual segmentations.
AB - Video object segmentation (VOS) is important for various computer vision problems, and handling it with minimal human supervision is highly desired for the large-scale applications. To bring down the supervision, existing approaches largely follow a data mining perspective by assuming the availability of multiple videos sharing the same object categories. It, however, would be problematic for the tasks that consume a single video. To address this problem, this paper proposes a novel approach that explores weakly labeled images to solve video object segmentation. Given a video labeled with a target category, images labeled with the same category are collected, from which noisy object exemplars are automatically discovered. After that the proposed approach extracts a set of region proposals on various frames and efficiently matches them with massive noisy exemplars in terms of appearance and spatial context. We then jointly select the best proposals across the video by solving a novel submodular problem that combines region voting and global region matching. Finally, the localization results are leveraged as strong supervision to guide pixel-level segmentation. Extensive experiments are conducted on two challenging public databases: Youtube-Objects and DAVIS. The results suggest that the proposed approach improves over previous weakly supervised/unsupervised approaches significantly, showing a performance even comparable with the several approaches supervised by the costly manual segmentations.
KW - Semantic object segmentation
KW - exemplar matching
KW - submodular optimization
KW - weakly labeled video
UR - https://www.scopus.com/pages/publications/85042174494
U2 - 10.1109/TIP.2018.2806995
DO - 10.1109/TIP.2018.2806995
M3 - 文章
C2 - 29870345
AN - SCOPUS:85042174494
SN - 1057-7149
VL - 27
SP - 4245
EP - 4259
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
IS - 9
ER -